A lead-bismuth alloy smelting reaction stage intelligent judgment method based on image recognition
Patent Information
- Application Number
- CN202610381796.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-26
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2046-03-26
AI Technical Summary
这种人工判断方法存在显著缺陷:首先,判断结果严重依赖操作人员的个人经验水平,不同操作人员的判断标准存在差异,导致判断结果缺乏一致性和客观性;其次,人工观察受限于人眼的感知能力,难以捕捉熔体表面的微小变化和动态演化特征,容易遗漏关键的阶段转变信号;再次,高温熔炼环境恶劣,长时间连续观察容易造成操作人员疲劳,影响判断准确性;最后,人工判断方法无法实现实时记录和数据积累,不利于工艺优化和质量追溯
本发明通过构建全面的多维视觉特征提取体系,不仅提取常规的颜色直方图特征、纹理特征、光流特征和形态学特征,更重要的是针对铅铋合金熔炼过程的物理化学特性,创新性地引入渣层动态演化特征提取模块,包括渣层厚度变化率检测、界面波动频谱分析和渣层破裂模式识别三个维度。渣层作为铅铋合金熔炼过程中杂质富集和反应进行的重要界面,其动态演化直接反映反应阶段的进展状态。
Smart Images

Figure CN122289186B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of lead-bismuth alloy smelting control, and in particular to an intelligent judgment method for the reaction stage of lead-bismuth alloy smelting based on image recognition. Background Technology
[0002] Lead-bismuth alloys, as a crucial coolant material in liquid metal-cooled fast reactors of Generation IV advanced nuclear energy systems, possess excellent neutron physics properties, a high boiling point, and superior thermal conductivity, making them promising candidates for applications in the nuclear energy field. The quality of lead-bismuth alloy preparation directly impacts the reactor's operational safety and economics, with the smelting process being a key factor in ensuring its quality. During lead-bismuth alloy smelting, the material undergoes multiple reaction stages, including initial heating, melting, impurity removal, homogenization, and cooling solidification. The control of process parameters at each stage directly determines the purity, homogeneity, and physical properties of the final product. Therefore, accurately determining the stage of the smelting reaction and adjusting process parameters in a timely manner is of great significance for improving lead-bismuth alloy product quality, reducing production costs, and ensuring production safety.
[0003] Traditional methods for judging the reaction stage in lead-bismuth alloy smelting primarily rely on manual observation and experience. Operators judge the current reaction stage by observing superficial characteristics such as color changes on the melt surface, slag layer distribution, and bubble formation, combined with process parameters such as temperature and time, based on long-accumulated experience. This manual judgment method has significant drawbacks: First, the judgment results heavily depend on the operator's individual experience level, and different operators have different judgment standards, leading to a lack of consistency and objectivity in the results; second, manual observation is limited by the perception of the human eye, making it difficult to capture subtle changes and dynamic evolution characteristics of the melt surface, easily missing key stage transition signals; third, the harsh high-temperature smelting environment and prolonged continuous observation can easily cause operator fatigue, affecting the accuracy of judgment; finally, the manual judgment method cannot achieve real-time recording and data accumulation, which is not conducive to process optimization and quality traceability.
[0004] With the development of automation technology, some enterprises have begun to adopt automatic judgment methods based on temperature and time parameters. This method presets the temperature range and duration for different reaction stages, and determines the corresponding stage when the actual temperature and time meet the preset conditions. However, this method only considers two single-dimensional parameters—temperature and time—ignoring the rich physicochemical information during the smelting process, such as melt surface morphology, slag evolution, and bubble generation. The actual smelting process is affected by various factors, including fluctuations in raw material composition, changes in furnace conditions, and environmental factors. Under the same temperature and time conditions, different reaction states may correspond, resulting in low accuracy and a high risk of misjudgment based on a single parameter. Furthermore, this method cannot provide early warning of stage transitions and lacks the ability to analyze the dynamic evolution trend of the smelting process, making it difficult to meet the needs of refined process control.
[0005] In recent years, some studies have attempted to apply image recognition technology to the monitoring of the smelting process, identifying stages by acquiring images of the melt surface and extracting simple color or texture features. However, existing methods generally suffer from the following shortcomings: First, the feature extraction dimension is singular, focusing only on static color or texture information and failing to fully explore the time-varying features unique to the smelting process, such as the dynamic evolution of the slag layer, interface fluctuations, and bubble rupture. Second, there is a lack of deep modeling of multi-scale time-series information, failing to effectively capture rapid changes within the reaction stage and the slow evolutionary trends of stage transitions. Third, the judgment model structure is simple, often employing traditional machine learning methods or shallow neural networks, which are insufficient for representing complex nonlinear spatiotemporal features. Fourth, there is a lack of early warning mechanisms for stage transitions, allowing only post-event judgment of the current stage and failing to identify impending stage transitions in advance, thus limiting the response time for process adjustments. These shortcomings result in low accuracy and poor robustness of existing image recognition methods in practical applications, making it difficult to meet the stringent requirements of judgment accuracy and real-time performance in the industrial production of lead-bismuth alloy smelting.
[0006] Therefore, there is an urgent need to develop an intelligent judgment method for the reaction stage of lead-bismuth alloy smelting that comprehensively utilizes multi-dimensional visual features, deeply mines the spatiotemporal evolution law, and has high accuracy and early warning capabilities, in order to overcome the shortcomings of existing technologies and promote the intelligent upgrading of lead-bismuth alloy smelting processes. Summary of the Invention
[0007] In view of this, the present invention provides an intelligent judgment method for the reaction stage of lead-bismuth alloy smelting based on image recognition. The purpose is to construct a digital twin model that is synchronized with the construction site in real time, and combine graph neural network technology to model and analyze the correlation and transmission mechanism between risk factors. This establishes a multi-level assessment system that integrates component-level and process-level risks, enabling real-time, dynamic, comprehensive, and accurate risk assessment of the steel-wood structure construction process. This provides a scientific basis and technical support for safety management and decision-making at the construction site, effectively reducing the probability of construction safety accidents and ensuring the safety of construction personnel and the quality of the project.
[0008] To achieve the above objectives, this invention provides an intelligent judgment method for the reaction stage of lead-bismuth alloy smelting based on image recognition, comprising the following steps: B1: Multidimensional visual feature extraction is performed on the acquired lead-bismuth alloy melt surface image sequence, extracting color histogram features, slag layer texture features, optical flow features of flow morphology, and morphological features of bubble rupture. Furthermore, the dynamic evolution features of the slag-melt interface are extracted from the acquired lead-bismuth alloy melt surface image sequence. Specifically, through slag layer thickness change rate detection, interface fluctuation spectrum analysis, and slag layer rupture mode recognition, the unique dynamic evolution information of the slag layer during lead-bismuth alloy smelting is captured. Various features are then combined to output an enhanced multidimensional feature vector sequence. B2: Receives the enhanced multidimensional feature vector sequence, uses principal component analysis to reduce the dimensionality of the enhanced multidimensional feature vector sequence, and uses dynamic time warping algorithm to align the time axis of the dimensionality-reduced time series data, outputting the aligned low-dimensional time series feature sequence; B3: Receive the aligned low-dimensional temporal feature sequence and input it into the cascaded model based on multi-scale spatiotemporal feature fusion. The cascaded model of multi-scale spatiotemporal feature fusion uses gradient boosting tree to extract key short-term spatial features, uses dual-path long short-term memory network to extract fast-changing and slow-evolving temporal patterns respectively, dynamically assigns weights to spatiotemporal features through self-attention mechanism, and integrates spatiotemporal features of different scales through adaptive feature fusion mechanism to output multi-scale dynamic spatiotemporal feature representation; B4: Receives multi-scale dynamic spatiotemporal feature representations, inputs them into a fully connected neural network layer for nonlinear mapping, performs normalization processing through a multi-classification activation function, and outputs probability distribution sequences corresponding to each lead-bismuth alloy smelting reaction stage; further constructs a reaction stage transition early warning mechanism to identify the moment when the reaction stage is about to change in advance and outputs a reaction stage transition early warning signal. B5: Receive the probability distribution sequence of each lead-bismuth alloy smelting reaction stage, use a sliding window mechanism to smooth the probability distribution within consecutive time steps, select the category corresponding to the maximum probability after smoothing, and output the judgment result of the lead-bismuth alloy smelting reaction stage.
[0009] As a further improvement of the present invention: Optionally, step B1 further includes: A high-temperature industrial camera was used to acquire a sequence of images of the surface of the melt during the smelting process of lead-bismuth alloy, thereby obtaining continuous images of the melt surface. The image sequence contains surface feature information of different reaction stages during the smelting process. Multidimensional visual feature extraction is performed on the acquired image sequence, including color histogram feature extraction, texture feature extraction, optical flow feature extraction, and morphological feature extraction. The color histogram feature extraction involves statistically analyzing the pixel distribution of each color channel in the image within the color space, converting the image to the HSV color space, and calculating histograms for the hue, saturation, and brightness channels. Each channel's histogram is divided into 32 bins, resulting in a color histogram feature vector. superscript Indicates the time index; The texture feature extraction is achieved by analyzing the gray-level co-occurrence matrix (GLCM) of the image and calculating four texture description parameters: contrast, correlation, energy, and entropy. The calculation directions of the GLCM are set to 0 degrees, 45 degrees, 90 degrees, and 135 degrees, and the pixel-to-pixel distance is set to 1 pixel, resulting in a texture feature vector. ; The optical flow feature extraction is achieved by calculating the pixel motion vector field between adjacent frames. The Farneback dense optical flow estimation algorithm is used to calculate the optical flow field between adjacent frames in the image sequence. The average amplitude, principal direction, and directional dispersion of the optical flow field are extracted as optical flow features to obtain the optical flow feature vector. ; The morphological feature extraction involves performing morphological processing on the image, using morphological opening and closing operations to extract bubble regions. The morphological operation's structure element is set to a 5×5 circular kernel, calculating the number, average area, and distribution density of bubbles to obtain the morphological feature vector. ; The dynamic evolution features of the slag-melt interface are extracted from the image sequence of the melt surface during the lead-bismuth alloy smelting process. The dynamic evolution features of the slag-melt interface include slag thickness change rate detection, interface fluctuation spectrum analysis and slag fracture mode identification. The slag layer thickness change rate detection method identifies the boundary between the slag layer and the melt using an edge detection algorithm. The Canny edge detection algorithm is used for edge extraction of the image, with dual thresholds set at 20% and 40% of the pixel grayscale range. The boundary line position is fitted using Hough transform, and the Euclidean distance difference between adjacent frames is calculated as the slag layer thickness change. The slag layer thickness change rate is obtained by dividing the slag layer thickness change by the time interval. ; The interface wave spectrum analysis extracts the dominant frequency components and spectral energy distribution by performing a Fast Fourier Transform on the boundary position sequence, characterizing the periodicity of interface waves. Specifically, for continuous... The sequence of boundary positions at each moment Perform a Fast Fourier Transform to obtain the frequency domain representation. Before extraction The frequency components corresponding to the maximum amplitude are taken as the dominant frequencies. The total spectral energy and the proportion of dominant frequency energy are calculated to obtain the spectral feature vector of the interface fluctuation. ; The slag layer fracture pattern recognition identifies the void regions formed by slag layer fracture through connected component analysis. First, the image is binarized, and a threshold of 0.8 times the average pixel grayscale value is set. Morphological closing operations are then used to fill small voids, with the structuring element of the morphological closing operation set to a 5×5 circular kernel. Void regions are extracted through connected component analysis, and the number of voids is calculated. Total area of holes and the spatial distribution entropy of holes The spatial distribution entropy is obtained by dividing the image into × By analyzing the grid, the percentage of voids within each grid cell is statistically analyzed, and the information entropy is calculated to obtain the feature vector of the slag layer fracture mode. ; The slag layer thickness change rate, the interfacial fluctuation spectrum feature vector, and the slag layer fracture mode feature vector are concatenated and combined to obtain the time step. slag layer dynamic evolution feature vector The calculation formula is: ; The color histogram feature vector, texture feature vector, optical flow feature vector, morphological feature vector, and slag layer dynamic evolution feature vector are concatenated and combined to obtain the time step. Enhanced multidimensional feature vectors The calculation formula is: ; in, This represents a vector concatenation operation; Multidimensional visual features and dynamic evolution features of the slag-melt interface are extracted from each frame of the entire image sequence and combined to obtain an enhanced multidimensional feature vector sequence.
[0010] Optionally, step B2 further includes: Principal component analysis is used to reduce the dimensionality of the enhanced multidimensional feature vector sequence. The principal component analysis method achieves data dimensionality reduction by calculating the covariance matrix of the feature vectors and extracting the feature vectors corresponding to their main eigenvalues. Calculate the covariance matrix of the enhanced multidimensional eigenvector sequence, perform eigenvalue decomposition on the covariance matrix, and select the first... The eigenvectors corresponding to the largest eigenvalues form a projection matrix, which projects the enhanced multidimensional eigenvectors onto the principal component space to obtain a sequence of low-dimensional eigenvectors after dimensionality reduction. The time axis is aligned by a dynamic time warping algorithm on the dimensionality-reduced low-dimensional feature vector sequence. A reference time-series feature sequence is defined, and the cumulative distance matrix between the current dimensionality-reduced low-dimensional feature vector sequence and the reference time-series feature sequence is calculated. The cumulative distance matrix uses Euclidean distance as the distance metric between the two feature vectors. The optimal time alignment path is found by dynamic programming to obtain the aligned low-dimensional time-series feature sequence.
[0011] Optionally, step B3 further includes: A cascaded model based on multi-scale spatiotemporal feature fusion is constructed, which includes a gradient boosting tree module, a dual-path long short-term memory network module, a self-attention mechanism module, and an adaptive feature fusion module. The gradient boosting tree module receives the feature vector of each time step from the aligned low-dimensional temporal feature sequence, extracts nonlinear spatial key features using the gradient boosting tree algorithm, and trains each decision tree by fitting the residual of the previous decision tree, outputting the time step. Spatial key feature vector sequence ; The dual-path long short-term memory network module includes a fast path and a slow path, which process the spatial key feature vector sequence in parallel. The fast path employs a long short-term memory network with a small time window to capture rapidly changing features within the reaction phase. The time window size of the fast path is set to 5 time steps. The long short-term memory network of the fast path contains one long short-term memory unit layer, and the hidden state dimension of each long short-term memory unit is set to 128. The output time... Fast path hidden state sequence ; The slow path employs a long short-term memory (LSTM) network with a large time window to capture the slow evolutionary trend of reaction phase transitions. The time window size of the slow path is set to 20 time steps. The LSTM network of the slow path contains two layers of LSTM units, with the hidden state dimension of each LSTM unit set to 64. The output time... Slow path hidden state sequence ; The fast path hidden state sequence and the slow path hidden state sequence are concatenated to obtain the time step. Multipath Feature Sequence The calculation formula is: ; The self-attention mechanism module dynamically assigns weights to the multi-path feature sequence. First, it transforms the multi-path feature sequence into a query vector, a key vector, and a value vector through three independent linear transformation matrices. The dimension of the linear transformation matrix is set to the multi-path feature dimension × the attention head dimension. The attention score is obtained by calculating the dot product of the query vector and the key vector. The attention score is normalized using Softmax to obtain the attention weight, and the formula for calculating the attention weight is as follows: ; in, Indicates the first At the nth moment, for the first Normalized attention weights at each moment, Indicates the first At the nth moment, for the first Attention score at each moment, Represents an exponential function. This represents the length of the aligned low-dimensional temporal feature sequence. For summation index; The attention-enhanced feature representation is obtained by weighting and summing the attention weights and the value vector. The calculation formula is: ; in, Indicates the first The attention enhancement features at each time point are represented as follows: Indicates the first A vector of values at each moment; The adaptive feature fusion module dynamically adjusts the weights of multi-path features and attention-enhanced features through a gating mechanism, thus integrating the multi-path features... and attention enhancement features The data is concatenated and input into a gating network, which consists of two fully connected neural networks. The first fully connected neural network has 0.5 times the number of neurons in the concatenation feature dimension and uses the ReLU activation function. The second fully connected neural network has the same number of neurons as the concatenation feature dimension and uses the Sigmoid activation function. The output is the first fully connected neural network. The fusion gate weight vector at each time step ; Multi-scale dynamic spatiotemporal feature representation is calculated through weighted fusion, and the calculation formula is as follows: ; in, Indicates the first Multi-scale dynamic spatiotemporal feature representation at each moment This represents element-wise multiplication; The entire aligned low-dimensional temporal feature sequence is input into a cascaded model based on multi-scale spatiotemporal feature fusion to obtain a multi-scale dynamic spatiotemporal feature representation sequence.
[0012] Optionally, step B4 further includes: A fully connected neural network layer is constructed to perform nonlinear mapping and classification decisions on multi-scale dynamic spatiotemporal features. The fully connected neural network layer includes a hidden layer and an output layer. The hidden layer receives multi-scale dynamic spatiotemporal feature representations and performs feature transformation through linear transformation and ReLU nonlinear activation function; The output layer receives the output of the hidden layer and obtains the original scores of each reaction stage category through linear transformation. The number of neurons in the output layer is set to the number of categories in the lead-bismuth alloy smelting reaction stage. The original scores were normalized using the Softmax multi-class activation function to obtain the probability distribution of each lead-bismuth alloy smelting reaction stage. The entire multi-scale dynamic spatiotemporal feature representation sequence is input into the fully connected neural network layer and normalized to obtain the probability distribution sequence of each lead-bismuth alloy smelting reaction stage. Based on the obtained probability distribution sequence of each lead-bismuth alloy smelting reaction stage, a reaction stage transition early warning mechanism is constructed, which includes KL divergence detection. The KL divergence detection measures the degree of change in the probability distribution by calculating the KL divergence between the probability distribution at the current time and the previous time. The formula for calculating the KL divergence is as follows: ; in, Indicates the first The moment and the first KL divergence between probability distributions at time points This indicates the number of categories in the lead-bismuth alloy smelting reaction stage. Indicates the first The time belongs to the first The probability of each type of reaction stage in the smelting of a lead-bismuth alloy, when the KL divergence Exceeding the preset threshold Time-based change warning; When KL divergence detection triggers a transition warning, a reaction stage transition warning signal is output. The advance time of the reaction stage transition warning is 2 to 5 time steps, providing a buffer time for process adjustment. The training process for the cascaded model and the fully connected neural network layer includes: The first stage is training the gradient boosting tree module: a training dataset containing labeled reaction stage categories is constructed, the low-dimensional time-series feature sequences in the training dataset are input into the gradient boosting tree module, and the gradient boosting tree algorithm is used for training. Each decision tree is iteratively optimized by fitting the residual of the previous decision tree. After training is completed, the parameters of the gradient boosting tree module are fixed. In the second stage, the neural network modules and fully connected neural network layers are jointly trained: the low-dimensional temporal feature sequences from the training dataset are input into the gradient boosting tree module with fixed parameters to obtain the spatial key feature vector sequence. The spatial key feature vector sequence is then input into the dual-path long short-term memory network module, the self-attention mechanism module, and the adaptive feature fusion module in sequence to obtain the multi-scale dynamic spatiotemporal feature representation. The multi-scale dynamic spatiotemporal feature representation is then input into the fully connected neural network layer to obtain the predicted reaction stage probability distribution. The cross-entropy loss between the predicted reaction stage probability distribution and the actual reaction stage label is calculated. The backpropagation algorithm is used to calculate the gradient of the loss function with respect to the parameters of the dual-path long short-term memory network module, the self-attention mechanism module, the adaptive feature fusion module, and the fully connected neural network layer. The Adam optimization algorithm is used to update the parameters of all modules synchronously. The training process is repeated until the cross-entropy loss converges or the preset maximum number of training rounds is reached.
[0013] Optionally, step B5 further includes: A sliding window mechanism is used to smooth the probability distribution within consecutive time steps; Set the sliding window length. For each time step, take the time steps before and after that time step to form the sliding window. Calculate the average value of the probability distribution at each time step within the sliding window to obtain the smoothed probability distribution. In the smoothed probability distribution, the category corresponding to the maximum probability is selected as the judgment result of the lead-bismuth alloy smelting reaction stage at that moment; Repeat the sliding window smoothing and category selection process for the entire probability distribution sequence to output a complete sequence of judgment results for the lead-bismuth alloy smelting reaction stage.
[0014] Compared with the prior art, the present invention has at least the following beneficial effects: This invention constructs a comprehensive multi-dimensional visual feature extraction system, which not only extracts conventional color histogram features, texture features, optical flow features, and morphological features, but more importantly, it innovatively introduces a slag layer dynamic evolution feature extraction module targeting the physicochemical characteristics of the lead-bismuth alloy smelting process. This module includes three dimensions: slag layer thickness change rate detection, interface fluctuation spectrum analysis, and slag layer fracture mode recognition. As a crucial interface for impurity enrichment and reaction during the lead-bismuth alloy smelting process, the dynamic evolution of the slag layer directly reflects the progress of the reaction stages.
[0015] This invention employs a cascaded model architecture based on multi-scale spatiotemporal feature fusion to achieve collaborative modeling of features at different time scales during lead-bismuth alloy smelting. The cascaded model first utilizes the gradient boosting tree algorithm to extract key features in the nonlinear space, fully leveraging the advantages of gradient boosting trees in handling high-dimensional nonlinear relationships. Then, a dual-path long short-term memory network is used to capture rapidly changing features and slowly evolving trends, respectively. The fast path uses a small time window to focus on instantaneous changes within a stage, while the slow path uses a large time window to focus on cross-stage evolutionary trends. Parallel processing of the two paths achieves comprehensive capture of information across multiple time scales. Furthermore, a self-attention mechanism is used to dynamically assign weights to spatiotemporal features, automatically focusing on important features at key moments. Finally, a gating mechanism is used to achieve adaptive feature fusion, dynamically adjusting the contribution weights of features at different scales.
[0016] This invention innovatively constructs a reaction stage transition early warning mechanism. It measures the degree of distribution change by calculating the KL divergence between probability distributions at consecutive time points. When the KL divergence exceeds a preset threshold, a transition early warning signal is triggered in advance, with an advance warning time of 2 to 5 time steps. This early warning capability provides valuable buffer time for operators or automatic control systems, allowing process parameter adjustments to begin before the actual stage transition occurs. This avoids the adjustment lag problem caused by traditional post-event judgment methods, greatly improving the precision and response speed of process control. Furthermore, this invention uses a sliding window mechanism to smooth and filter the probability distribution sequence, effectively suppressing the interference of single-frame image noise and random fluctuations on the judgment results, resulting in more stable and reliable output judgment results. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating an embodiment of the intelligent judgment method for the reaction stage of lead-bismuth alloy smelting based on image recognition according to an embodiment of the present invention. Figure 2 This is a schematic diagram illustrating the early warning effect of the reaction phase transition. Detailed Implementation
[0018] The present invention will be further described below with reference to the accompanying drawings, but this is not intended to limit the present invention in any way. Any modifications or substitutions made based on the teachings of the present invention shall fall within the protection scope of the present invention.
[0019] Example 1: An intelligent judgment method for the reaction stage of lead-bismuth alloy smelting based on image recognition, such as... Figure 1 As shown, it includes the following steps: B1: Multidimensional visual feature extraction is performed on the acquired lead-bismuth alloy melt surface image sequence, extracting color histogram features, slag layer texture features, optical flow features of flow morphology, and morphological features of bubble bursting. Furthermore, the dynamic evolution features of the slag-melt interface are extracted from the acquired lead-bismuth alloy melt surface image sequence, specifically through slag layer thickness change rate detection, interface fluctuation spectrum analysis, and slag layer fracture mode recognition to capture the unique dynamic evolution information of the slag layer during lead-bismuth alloy smelting. These features are then concatenated and combined to output an enhanced multidimensional feature vector sequence, including: A high-temperature industrial camera was used to acquire a sequence of images of the surface of the melt during the smelting process of lead-bismuth alloy, thereby obtaining continuous images of the melt surface. The image sequence contains surface feature information of different reaction stages during the smelting process. Multidimensional visual feature extraction is performed on the acquired image sequence, including color histogram feature extraction, texture feature extraction, optical flow feature extraction, and morphological feature extraction. The color histogram feature extraction involves statistically analyzing the pixel distribution of each color channel in the image within the color space, converting the image to the HSV color space, and calculating histograms for the hue, saturation, and brightness channels. Each channel's histogram is divided into 32 bins, resulting in a color histogram feature vector. superscript Indicates the time index; in this embodiment, the RGB to HSV color space conversion uses a standard conversion formula. For the H channel, it will... The range is evenly divided into 32 bins; for the S and V channels, The image was evenly divided into 32 bins; the frequency of each pixel in each channel and bin was counted, and after normalization, a 96-dimensional color histogram feature vector was obtained. ; The texture feature extraction is achieved by analyzing the gray-level co-occurrence matrix (GLCM) of the image and calculating four texture description parameters: contrast, correlation, energy, and entropy. The calculation directions of the GLCM are set to 0 degrees, 45 degrees, 90 degrees, and 135 degrees, and the pixel-to-pixel distance is set to 1 pixel, resulting in a texture feature vector. ; The optical flow feature extraction is achieved by calculating the pixel motion vector field between adjacent frames. The Farneback dense optical flow estimation algorithm is used to calculate the optical flow field between adjacent frames in the image sequence. The average amplitude, principal direction, and directional dispersion of the optical flow field are extracted as optical flow features to obtain the optical flow feature vector. In this embodiment, the Farneback algorithm employs a pyramid multi-scale strategy, constructing a 3-layer image pyramid with a polynomial expansion window size of 5×5 pixels and undergoing 3 iterations. It extracts the average amplitude, principal direction, and directional dispersion of the optical flow across the entire image to obtain a 3D optical flow feature vector. ; The morphological feature extraction involves performing morphological processing on the image, using morphological opening and closing operations to extract bubble regions. The morphological operation's structure element is set to a 5×5 circular kernel, calculating the number, average area, and distribution density of bubbles to obtain the morphological feature vector. In this embodiment, the Otsu thresholding method is used to binarize the grayscale image, and then opening and closing operations are performed to remove noise. An 8-connected component labeling algorithm is used to identify each independent bubble region, and the number, average area, and distribution density of bubbles are counted to obtain a 3D morphological feature vector. ; The dynamic evolution features of the slag-melt interface are extracted from the image sequence of the melt surface during the lead-bismuth alloy smelting process. The dynamic evolution features of the slag-melt interface include slag thickness change rate detection, interface fluctuation spectrum analysis and slag fracture mode identification. The slag layer thickness change rate detection method identifies the boundary between the slag layer and the melt using an edge detection algorithm. The Canny edge detection algorithm is used for edge extraction of the image, with dual thresholds set at 20% and 40% of the pixel grayscale range. The boundary line position is fitted using Hough transform, and the Euclidean distance difference between adjacent frames is calculated as the slag layer thickness change. The slag layer thickness change rate is obtained by dividing the slag layer thickness change by the time interval. In this embodiment, Canny edge detection uses a 5×5 Gaussian filter for smoothing, with the standard deviation of the Gaussian kernel set to 1.4, and then the Sobel operator is used to calculate the image gradient; in the Hough linear transform... The quantization step size is set to 1 pixel. The quantization step size is set to 1 degree, and the voting threshold is set to 10% of the total number of edge points. The ordinate of the boundary line at the horizontal center of the image is taken as the boundary line position feature. The position difference between adjacent frames is calculated and divided by the time interval of 0.1 seconds to obtain the slag layer thickness change rate. ; The interface wave spectrum analysis extracts the dominant frequency components and spectral energy distribution by performing a Fast Fourier Transform on the boundary position sequence, characterizing the periodicity of interface waves. Specifically, for continuous... The sequence of boundary positions at each moment Perform a Fast Fourier Transform to obtain the frequency domain representation. Before extraction The frequency components corresponding to the maximum amplitude are taken as the dominant frequencies. The total spectral energy and the proportion of dominant frequency energy are calculated to obtain the spectral feature vector of the interface fluctuation. In this embodiment, the sliding window length is set. Before extraction The frequency corresponding to the maximum amplitude value is taken as the main frequency; The slag layer fracture pattern recognition identifies the void regions formed by slag layer fracture through connected component analysis. First, the image is binarized, and a threshold of 0.8 times the average pixel grayscale value is set. Morphological closing operations are then used to fill small voids, with the structuring element of the morphological closing operation set to a 5×5 circular kernel. Void regions are extracted through connected component analysis, and the number of voids is calculated. Total area of holes and the spatial distribution entropy of holes The spatial distribution entropy is obtained by dividing the image into × By analyzing the grid and calculating the percentage of voids within each grid cell, the information entropy is obtained, thus yielding the feature vector of the slag layer fracture mode. In this embodiment, the binarization threshold is set to... ,in The image is represented by the pixel mean of the grayscale image; an 8-connected connected component labeling algorithm is used to identify hole regions; the image is divided into... The grid consists of 36 grid cells, used to calculate the spatial distribution entropy of the holes. ,in For the first The percentage of holes within each grid ; Obtain the 3D slag layer fracture mode feature vector ; The slag layer thickness change rate, the interfacial fluctuation spectrum feature vector, and the slag layer fracture mode feature vector are concatenated and combined to obtain the time step. slag layer dynamic evolution feature vector The calculation formula is: ; The color histogram feature vector, texture feature vector, optical flow feature vector, morphological feature vector, and slag layer dynamic evolution feature vector are concatenated and combined to obtain the time step. Enhanced multidimensional feature vectors The calculation formula is: ; in, This represents a vector concatenation operation; Multidimensional visual features and dynamic evolution features of the slag-melt interface are extracted from each frame of the entire image sequence and combined to obtain an enhanced multidimensional feature vector sequence.
[0020] B2: Receives the enhanced multidimensional feature vector sequence, performs dimensionality reduction on the enhanced multidimensional feature vector sequence using principal component analysis, and aligns the time-series data of the dimensionality reduction using a dynamic time warping algorithm, outputting the aligned low-dimensional time-series feature sequence, including: Principal component analysis is used to reduce the dimensionality of the enhanced multidimensional feature vector sequence. The principal component analysis method achieves data dimensionality reduction by calculating the covariance matrix of the feature vectors and extracting the feature vectors corresponding to their main eigenvalues. Calculate the covariance matrix of the enhanced multidimensional eigenvector sequence, perform eigenvalue decomposition on the covariance matrix, and select the first... The eigenvectors corresponding to the largest eigenvalues form a projection matrix, which projects the enhanced multidimensional eigenvectors onto the principal component space, resulting in a dimensionality-reduced low-dimensional eigenvector sequence. In this embodiment, after centering the eigenma matrix, the covariance matrix is calculated and eigenvalue decomposition is performed. The top eigenvalues are then selected. The eigenvectors corresponding to the largest eigenvalues constitute the projection matrix; The reduced-dimensional feature vector sequence is aligned on the time axis using a dynamic time warping algorithm. A reference time-series feature sequence is defined, and the cumulative distance matrix between the current reduced-dimensional feature vector sequence and the reference time-series feature sequence is calculated. The cumulative distance matrix uses Euclidean distance as the distance metric between the two feature vectors. The optimal time alignment path is found through dynamic programming to obtain the aligned low-dimensional time-series feature sequence. In this embodiment, the reference time-series feature sequence is a representative sequence selected from historical normal smelting batches. B3: Receives the aligned low-dimensional temporal feature sequence and inputs it into a cascaded model based on multi-scale spatiotemporal feature fusion. This cascaded model uses a gradient boosting tree to extract key short-term spatial features, employs a dual-path long short-term memory network to extract rapidly changing and slowly evolving temporal patterns respectively, dynamically assigns weights to spatiotemporal features through a self-attention mechanism, and then integrates spatiotemporal features of different scales through an adaptive feature fusion mechanism, outputting a multi-scale dynamic spatiotemporal feature representation, including: A cascaded model based on multi-scale spatiotemporal feature fusion is constructed, which includes a gradient boosting tree module, a dual-path long short-term memory network module, a self-attention mechanism module, and an adaptive feature fusion module. The gradient boosting tree module receives the feature vector of each time step from the aligned low-dimensional temporal feature sequence, extracts nonlinear spatial key features using the gradient boosting tree algorithm, and trains each decision tree by fitting the residual of the previous decision tree, outputting the time step. Spatial key feature vector sequence In this embodiment, the gradient boosting tree is implemented using XGBoost, with 300 decision trees, a maximum depth of 6 for each tree, and a learning rate of... Regularization parameters and The leaf node embedding method is used, with the embedding dimension set to 64, to obtain a 64-dimensional spatial key feature vector. ; The dual-path long short-term memory network module includes a fast path and a slow path, which process the spatial key feature vector sequence in parallel. The fast path employs a long short-term memory network with a small time window to capture rapidly changing features within the reaction phase. The time window size of the fast path is set to 5 time steps. The long short-term memory network of the fast path contains one long short-term memory unit layer, and the hidden state dimension of each long short-term memory unit is set to 128. The output time... Fast path hidden state sequence In this embodiment, the calculation of the fast path LSTM follows the standard LSTM formula, including the forget gate, input gate, output gate, cell state update, and hidden state update. The slow path employs a long short-term memory (LSTM) network with a large time window to capture the slow evolutionary trend of reaction phase transitions. The time window size of the slow path is set to 20 time steps. The LSTM network of the slow path contains two layers of LSTM units, with the hidden state dimension of each LSTM unit set to 64. The output time... Slow path hidden state sequence In this embodiment, the slow path uses a two-layer stacked LSTM structure, with each layer having a hidden state dimension of 64. The first LSTM layer processes the spatial key feature sequence, and the second LSTM layer processes the output of the first layer, ultimately outputting a 64-dimensional slow path hidden state. ; The fast path hidden state sequence and the slow path hidden state sequence are concatenated to obtain the time step. Multipath Feature Sequence The calculation formula is: ; The self-attention mechanism module dynamically assigns weights to the multi-path feature sequence. First, it transforms the multi-path feature sequence into a query vector, a key vector, and a value vector through three independent linear transformation matrices. The dimension of the linear transformation matrix is set to the multi-path feature dimension × the attention head dimension; in this embodiment, the attention head dimension is set to 64. The attention score is obtained by calculating the dot product of the query vector and the key vector. The attention score is normalized using Softmax to obtain the attention weight, and the formula for calculating the attention weight is as follows: ; in, Indicates the first At the nth moment, for the first Normalized attention weights at each moment, Indicates the first At the nth moment, for the first Attention score at each moment, Represents an exponential function. This represents the length of the aligned low-dimensional temporal feature sequence. For summation index; The attention-enhanced feature representation is obtained by weighting and summing the attention weights and the value vector. The calculation formula is: ; in, Indicates the first The attention enhancement features at each time point are represented as follows: Indicates the first A vector of values at each moment; The adaptive feature fusion module dynamically adjusts the weights of multi-path features and attention-enhanced features through a gating mechanism, thus integrating the multi-path features... and attention enhancement features The data is concatenated and input into a gating network, which consists of two fully connected neural networks. The first fully connected neural network has 0.5 times the number of neurons in the concatenation feature dimension and uses the ReLU activation function. The second fully connected neural network has the same number of neurons as the concatenation feature dimension and uses the Sigmoid activation function. The output is the first fully connected neural network. The fusion gate weight vector at each time step In this embodiment, the first layer of the gating network has 128 neurons, and the second layer has 256 neurons. Multi-scale dynamic spatiotemporal feature representation is calculated through weighted fusion, and the calculation formula is as follows: ; in, Indicates the first Multi-scale dynamic spatiotemporal feature representation at each moment This represents element-wise multiplication; The entire aligned low-dimensional temporal feature sequence is input into a cascaded model based on multi-scale spatiotemporal feature fusion to obtain a multi-scale dynamic spatiotemporal feature representation sequence.
[0021] B4: Receives multi-scale dynamic spatiotemporal feature representations, inputs them into a fully connected neural network layer for nonlinear mapping, performs normalization through a multi-classification activation function, and outputs probability distribution sequences corresponding to each lead-bismuth alloy smelting reaction stage; further, it constructs a reaction stage transition early warning mechanism to identify the moment when a reaction stage transition is about to occur, and outputs reaction stage transition early warning signals, including: A fully connected neural network layer is constructed to perform nonlinear mapping and classification decisions on multi-scale dynamic spatiotemporal features. The fully connected neural network layer includes a hidden layer and an output layer. The hidden layer receives multi-scale dynamic spatiotemporal feature representations and performs feature transformation through linear transformation and ReLU nonlinear activation function; The output layer receives the output of the hidden layer and obtains the original scores of each reaction stage category through linear transformation. The number of neurons in the output layer is set to the number of categories of the lead-bismuth alloy smelting reaction stage. In this embodiment, the lead-bismuth alloy smelting reaction stage is divided into 5 categories: initial heating stage, melting stage, impurity removal stage, homogenization stage, and cooling and solidification stage. The original scores were normalized using the Softmax multi-class activation function to obtain the probability distribution of each lead-bismuth alloy smelting reaction stage. The entire multi-scale dynamic spatiotemporal feature representation sequence is input into the fully connected neural network layer and normalized to obtain the probability distribution sequence of each lead-bismuth alloy smelting reaction stage. Based on the obtained probability distribution sequence of each lead-bismuth alloy smelting reaction stage, a reaction stage transition early warning mechanism is constructed, which includes KL divergence detection. The KL divergence detection measures the degree of change in the probability distribution by calculating the KL divergence between the probability distribution at the current time and the previous time. The formula for calculating the KL divergence is as follows: ; in, Indicates the first The moment and the first KL divergence between probability distributions at time points This indicates the number of categories in the lead-bismuth alloy smelting reaction stage. Indicates the first The time belongs to the first The probability of each type of reaction stage in the smelting of a lead-bismuth alloy, when the KL divergence Exceeding the preset threshold A change warning is triggered at a certain time; in this embodiment, a preset threshold is used. To avoid zero values in logarithmic operations, a very small constant is added. ; When KL divergence detection triggers a transition warning, a reaction stage transition warning signal is output. The lead time for the reaction stage transition warning is 2 to 5 time steps, providing a buffer time for process adjustments. Figure 2 As shown; The training process for the cascaded model and the fully connected neural network layer includes: In the first stage, the gradient boosting tree module is trained: a training dataset containing labeled reaction stage categories is constructed. The low-dimensional time-series feature sequences in the training dataset are input into the gradient boosting tree module, and the gradient boosting tree algorithm is used for training. Each decision tree is iteratively optimized by fitting the residual of the previous decision tree. After training, the parameters of the gradient boosting tree module are fixed. In this embodiment, the training dataset contains data from 50 historical smelting batches, with 80% used for training and 20% for validation. GBDT training uses a multi-class cross-entropy loss function and an early stopping strategy to prevent overfitting. In the second stage, the neural network modules and fully connected neural network layers are jointly trained: Low-dimensional temporal feature sequences from the training dataset are input into a gradient boosting tree module with fixed parameters to obtain spatial key feature vector sequences. These sequences are then sequentially input into a dual-path long short-term memory network module, a self-attention mechanism module, and an adaptive feature fusion module to obtain multi-scale dynamic spatiotemporal feature representations. These multi-scale dynamic spatiotemporal feature representations are then input into the fully connected neural network layer to obtain the predicted reaction stage probability distribution. The cross-entropy loss between the predicted reaction stage probability distribution and the actual reaction stage labels is calculated. Backpropagation is used to calculate the gradient of the loss function with respect to the parameters of the dual-path long short-term memory network module, the self-attention mechanism module, the adaptive feature fusion module, and the fully connected neural network layer. The Adam optimization algorithm is used to synchronously update the parameters of all modules. This training process is repeated until the cross-entropy loss converges or the preset maximum number of training rounds is reached. In this embodiment, the batch size is set to 32; the Adam optimizer parameters are set as follows: learning rate... The maximum number of training rounds is 100, an early stop strategy is adopted, and the Dropout ratio is set to 0.3.
[0022] B5: Receives the probability distribution sequence for each lead-bismuth alloy smelting reaction stage, uses a sliding window mechanism to smooth the probability distribution within consecutive time steps, selects the category corresponding to the maximum probability after smoothing, and outputs the judgment result for the lead-bismuth alloy smelting reaction stage, including: A sliding window mechanism is used to smooth the probability distribution within consecutive time steps; The sliding window length is set. For each time step, the sliding window is formed by taking the time steps before and after that time step. The average value of the probability distribution at each time step within the sliding window is calculated to obtain the smoothed probability distribution. In this embodiment, the sliding window length is set to 9 time steps. In the smoothed probability distribution, the category corresponding to the maximum probability is selected as the judgment result of the lead-bismuth alloy smelting reaction stage at that moment; Repeat the sliding window smoothing and category selection process for the entire probability distribution sequence to output a complete sequence of judgment results for the lead-bismuth alloy smelting reaction stage.
[0023] It should be noted that the sequence numbers of the above embodiments of the present invention are merely for descriptive purposes and do not represent the superiority or inferiority of the embodiments. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, apparatus, article, or method. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.
[0024] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0025] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.
Claims
1. A method for intelligent judgment of the reaction stage in lead-bismuth alloy smelting based on image recognition, characterized in that, Includes the following steps: B1: Multidimensional visual feature extraction is performed on the acquired lead-bismuth alloy melt surface image sequence, extracting color histogram features, slag layer texture features, optical flow features of flow morphology, and morphological features of bubble rupture. Furthermore, the dynamic evolution features of the slag-melt interface are extracted from the acquired lead-bismuth alloy melt surface image sequence. Specifically, through slag layer thickness change rate detection, interface fluctuation spectrum analysis, and slag layer rupture mode recognition, the unique dynamic evolution information of the slag layer during lead-bismuth alloy smelting is captured. Various features are then combined to output an enhanced multidimensional feature vector sequence. B2: Receives the enhanced multidimensional feature vector sequence, uses principal component analysis to reduce the dimensionality of the enhanced multidimensional feature vector sequence, and uses dynamic time warping algorithm to align the time axis of the dimensionality-reduced time series data, outputting the aligned low-dimensional time series feature sequence; B3: Receive the aligned low-dimensional temporal feature sequence and input it into the cascaded model based on multi-scale spatiotemporal feature fusion. The cascaded model of multi-scale spatiotemporal feature fusion uses gradient boosting tree to extract key short-term spatial features, uses dual-path long short-term memory network to extract fast-changing and slow-evolving temporal patterns respectively, dynamically assigns weights to spatiotemporal features through self-attention mechanism, and integrates spatiotemporal features of different scales through adaptive feature fusion mechanism to output multi-scale dynamic spatiotemporal feature representation; B4: Receives multi-scale dynamic spatiotemporal feature representations, inputs them into a fully connected neural network layer for nonlinear mapping, performs normalization processing through a multi-classification activation function, and outputs probability distribution sequences corresponding to each lead-bismuth alloy smelting reaction stage; further constructs a reaction stage transition early warning mechanism to identify the moment when the reaction stage is about to change in advance and outputs a reaction stage transition early warning signal. B5: Receive the probability distribution sequence of each lead-bismuth alloy smelting reaction stage, use a sliding window mechanism to smooth the probability distribution within consecutive time steps, select the category corresponding to the maximum probability after smoothing, and output the judgment result of the lead-bismuth alloy smelting reaction stage.
2. The intelligent judgment method for the reaction stage of lead-bismuth alloy smelting based on image recognition according to claim 1, characterized in that, Step B1 includes: A high-temperature industrial camera was used to acquire a sequence of images of the surface of the melt during the smelting process of lead-bismuth alloy, thereby obtaining continuous images of the melt surface. The image sequence contains surface feature information of different reaction stages during the smelting process. Multidimensional visual feature extraction is performed on the acquired image sequence, including color histogram feature extraction, texture feature extraction, optical flow feature extraction, and morphological feature extraction. The color histogram feature extraction involves statistically analyzing the pixel distribution of each color channel in the image within the color space, converting the image to the HSV color space, and calculating histograms for the hue, saturation, and brightness channels. Each channel's histogram is divided into 32 bins, resulting in a color histogram feature vector. superscript Indicates the time index; The texture feature extraction is achieved by analyzing the gray-level co-occurrence matrix (GLCM) of the image and calculating four texture description parameters: contrast, correlation, energy, and entropy. The calculation directions of the GLCM are set to 0 degrees, 45 degrees, 90 degrees, and 135 degrees, and the pixel-to-pixel distance is set to 1 pixel, resulting in a texture feature vector. ; The optical flow feature extraction is achieved by calculating the pixel motion vector field between adjacent frames. The Farneback dense optical flow estimation algorithm is used to calculate the optical flow field between adjacent frames in the image sequence. The average amplitude, principal direction, and directional dispersion of the optical flow field are extracted as optical flow features to obtain the optical flow feature vector. ; The morphological feature extraction involves performing morphological processing on the image, using morphological opening and closing operations to extract bubble regions. The morphological operation's structure element is set to a 5×5 circular kernel, calculating the number, average area, and distribution density of bubbles to obtain the morphological feature vector. ; The dynamic evolution features of the slag-melt interface are extracted from the image sequence of the melt surface during the lead-bismuth alloy smelting process. The dynamic evolution features of the slag-melt interface include slag thickness change rate detection, interface fluctuation spectrum analysis and slag fracture mode identification. The slag layer thickness change rate detection method identifies the boundary between the slag layer and the melt using an edge detection algorithm. The Canny edge detection algorithm is used for edge extraction of the image, with dual thresholds set at 20% and 40% of the pixel grayscale range. The boundary line position is fitted using Hough transform, and the Euclidean distance difference between adjacent frames is calculated as the slag layer thickness change. The slag layer thickness change rate is obtained by dividing the slag layer thickness change by the time interval. ; The interface wave spectrum analysis extracts the dominant frequency components and spectral energy distribution by performing a Fast Fourier Transform on the boundary position sequence, characterizing the periodicity of interface waves. Specifically, for continuous... The sequence of boundary positions at each moment Perform a Fast Fourier Transform to obtain the frequency domain representation. Before extraction The frequency components corresponding to the maximum amplitude are taken as the dominant frequencies. The total spectral energy and the proportion of dominant frequency energy are calculated to obtain the spectral feature vector of the interface fluctuation. ; The slag layer fracture pattern recognition identifies the void regions formed by slag layer fracture through connected component analysis. First, the image is binarized, and a threshold of 0.8 times the average pixel grayscale value is set. Morphological closing operations are then used to fill small voids, with the structuring element of the morphological closing operation set to a 5×5 circular kernel. Void regions are extracted through connected component analysis, and the number of voids is calculated. Total area of holes and the spatial distribution entropy of holes The spatial distribution entropy is obtained by dividing the image into × By analyzing the grid, the percentage of voids within each grid is counted, and the information entropy is calculated to obtain the feature vector of the slag layer fracture mode. ; The slag layer thickness change rate, the interfacial fluctuation spectrum feature vector, and the slag layer fracture mode feature vector are concatenated and combined to obtain the time step. slag layer dynamic evolution feature vector The calculation formula is: ; The color histogram feature vector, texture feature vector, optical flow feature vector, morphological feature vector, and slag layer dynamic evolution feature vector are concatenated and combined to obtain the time step. Enhanced multidimensional feature vectors The calculation formula is: ; in, This represents a vector concatenation operation; Multidimensional visual features and dynamic evolution features of the slag-melt interface are extracted from each frame of the entire image sequence and combined to obtain an enhanced multidimensional feature vector sequence.
3. The intelligent judgment method for the reaction stage of lead-bismuth alloy smelting based on image recognition according to claim 2, characterized in that, Step B2 includes: Principal component analysis is used to reduce the dimensionality of the enhanced multidimensional feature vector sequence. The principal component analysis method achieves data dimensionality reduction by calculating the covariance matrix of the feature vectors and extracting the feature vectors corresponding to their main eigenvalues. Calculate the covariance matrix of the enhanced multidimensional eigenvector sequence, perform eigenvalue decomposition on the covariance matrix, and select the first... The eigenvectors corresponding to the largest eigenvalues form a projection matrix, which projects the enhanced multidimensional eigenvectors onto the principal component space to obtain a sequence of low-dimensional eigenvectors after dimensionality reduction. The time axis is aligned by a dynamic time warping algorithm on the dimensionality-reduced low-dimensional feature vector sequence. A reference time-series feature sequence is defined, and the cumulative distance matrix between the current dimensionality-reduced low-dimensional feature vector sequence and the reference time-series feature sequence is calculated. The cumulative distance matrix uses Euclidean distance as the distance metric between the two feature vectors. The optimal time alignment path is found by dynamic programming to obtain the aligned low-dimensional time-series feature sequence.
4. The intelligent judgment method for the reaction stage of lead-bismuth alloy smelting based on image recognition according to claim 3, characterized in that, Step B3 includes: A cascaded model based on multi-scale spatiotemporal feature fusion is constructed, which includes a gradient boosting tree module, a dual-path long short-term memory network module, a self-attention mechanism module, and an adaptive feature fusion module. The gradient boosting tree module receives the feature vector of each time step from the aligned low-dimensional temporal feature sequence, extracts nonlinear spatial key features using the gradient boosting tree algorithm, and trains each decision tree by fitting the residual of the previous decision tree, outputting the time step. Spatial key feature vector sequence ; The dual-path long short-term memory network module includes a fast path and a slow path, which process the spatial key feature vector sequence in parallel. The fast path employs a long short-term memory network with a small time window to capture rapidly changing features within the reaction phase. The time window size of the fast path is set to 5 time steps. The long short-term memory network of the fast path contains one long short-term memory unit layer, and the hidden state dimension of each long short-term memory unit is set to 128. The output time... Fast path hidden state sequence ; The slow path employs a long short-term memory (LSTM) network with a large time window to capture the slow evolutionary trend of reaction phase transitions. The time window size of the slow path is set to 20 time steps. The LSTM network of the slow path contains two layers of LSTM units, with the hidden state dimension of each LSTM unit set to 64. The output time... Slow path hidden state sequence ; The fast path hidden state sequence and the slow path hidden state sequence are concatenated to obtain the time step. Multipath Feature Sequence The calculation formula is: ; The self-attention mechanism module dynamically assigns weights to the multi-path feature sequence. First, it transforms the multi-path feature sequence into a query vector, a key vector, and a value vector through three independent linear transformation matrices. The dimension of the linear transformation matrix is set to the multi-path feature dimension × the attention head dimension. The attention score is obtained by calculating the dot product of the query vector and the key vector. The attention score is normalized using Softmax to obtain the attention weight, and the formula for calculating the attention weight is as follows: ; in, Indicates the first At the nth moment, for the first Normalized attention weights at each moment, Indicates the first At the nth moment, for the first Attention score at each moment, Represents an exponential function. This represents the length of the aligned low-dimensional temporal feature sequence. For summation index; The attention-enhanced feature representation is obtained by weighting and summing the attention weights and the value vector. The calculation formula is: ; in, Indicates the first The attention enhancement features at each time point are represented as follows: Indicates the first A vector of values at each moment; The adaptive feature fusion module dynamically adjusts the weights of multi-path features and attention-enhanced features through a gating mechanism, thus integrating the multi-path features... and attention enhancement features The data is concatenated and input into a gating network, which consists of two fully connected neural networks. The first fully connected neural network has 0.5 times the number of neurons in the concatenation feature dimension and uses the ReLU activation function. The second fully connected neural network has the same number of neurons as the concatenation feature dimension and uses the Sigmoid activation function. The output is the first fully connected neural network. The fusion gate weight vector at each time step ; Multi-scale dynamic spatiotemporal feature representation is calculated through weighted fusion, and the calculation formula is as follows: ; in, Indicates the first Multi-scale dynamic spatiotemporal feature representation at each moment This represents element-wise multiplication; The entire aligned low-dimensional temporal feature sequence is input into a cascaded model based on multi-scale spatiotemporal feature fusion to obtain a multi-scale dynamic spatiotemporal feature representation sequence.
5. The intelligent judgment method for the reaction stage of lead-bismuth alloy smelting based on image recognition according to claim 4, characterized in that, Step B4 includes: A fully connected neural network layer is constructed to perform nonlinear mapping and classification decisions on multi-scale dynamic spatiotemporal features. The fully connected neural network layer includes a hidden layer and an output layer. The hidden layer receives multi-scale dynamic spatiotemporal feature representations and performs feature transformation through linear transformation and ReLU nonlinear activation function; The output layer receives the output of the hidden layer and obtains the original scores of each reaction stage category through linear transformation. The number of neurons in the output layer is set to the number of categories in the lead-bismuth alloy smelting reaction stage. The original scores were normalized using the Softmax multi-class activation function to obtain the probability distribution of each lead-bismuth alloy smelting reaction stage. The entire multi-scale dynamic spatiotemporal feature representation sequence is input into the fully connected neural network layer and normalized to obtain the probability distribution sequence of each lead-bismuth alloy smelting reaction stage. Based on the obtained probability distribution sequence of each lead-bismuth alloy smelting reaction stage, a reaction stage transition early warning mechanism is constructed, which includes KL divergence detection. The KL divergence detection measures the degree of change in the probability distribution by calculating the KL divergence between the probability distribution at the current time and the previous time. The formula for calculating the KL divergence is as follows: ; in, Indicates the first The moment and the first KL divergence between probability distributions at time points This indicates the number of categories in the lead-bismuth alloy smelting reaction stage. Indicates the first The time belongs to the first The probability of each type of reaction stage in the smelting of a lead-bismuth alloy, when the KL divergence Exceeding the preset threshold Time-based change warning; When KL divergence detection triggers a transition warning, a reaction stage transition warning signal is output. The advance time of the reaction stage transition warning is 2 to 5 time steps, providing a buffer time for process adjustment. The training process for the cascaded model and the fully connected neural network layer includes: The first stage is training the gradient boosting tree module: a training dataset containing labeled reaction stage categories is constructed, the low-dimensional time-series feature sequences in the training dataset are input into the gradient boosting tree module, and the gradient boosting tree algorithm is used for training. Each decision tree is iteratively optimized by fitting the residual of the previous decision tree. After training is completed, the parameters of the gradient boosting tree module are fixed. In the second stage, the neural network modules and fully connected neural network layers are jointly trained: the low-dimensional temporal feature sequences from the training dataset are input into the gradient boosting tree module with fixed parameters to obtain the spatial key feature vector sequence. The spatial key feature vector sequence is then input into the dual-path long short-term memory network module, the self-attention mechanism module, and the adaptive feature fusion module in sequence to obtain the multi-scale dynamic spatiotemporal feature representation. The multi-scale dynamic spatiotemporal feature representation is then input into the fully connected neural network layer to obtain the predicted reaction stage probability distribution. The cross-entropy loss between the predicted reaction stage probability distribution and the actual reaction stage label is calculated. The backpropagation algorithm is used to calculate the gradient of the loss function with respect to the parameters of the dual-path long short-term memory network module, the self-attention mechanism module, the adaptive feature fusion module, and the fully connected neural network layer. The Adam optimization algorithm is used to update the parameters of all modules synchronously. The training process is repeated until the cross-entropy loss converges or the preset maximum number of training rounds is reached.
6. The intelligent judgment method for the reaction stage of lead-bismuth alloy smelting based on image recognition according to claim 5, characterized in that, Step B5 includes: A sliding window mechanism is used to smooth the probability distribution within consecutive time steps; Set the sliding window length. For each time step, take the time steps before and after that time step to form the sliding window. Calculate the average value of the probability distribution at each time step within the sliding window to obtain the smoothed probability distribution. In the smoothed probability distribution, the category corresponding to the maximum probability is selected as the judgment result of the lead-bismuth alloy smelting reaction stage at that moment; Repeat the sliding window smoothing and category selection process for the entire probability distribution sequence to output a complete sequence of judgment results for the lead-bismuth alloy smelting reaction stage.
Citation Information
Patent Citations
Pig feed fermentation process state analysis method based on image recognition
CN121685993A