Ball mill load identification method based on multi-modal signal fusion and attention enhancement
Through the multimodal signal fusion and attention enhancement method, combined with feature modal decomposition and symmetric point generation mode image technology, the accuracy and resource consumption problems of ball mill load state recognition are solved, and efficient load recognition and energy consumption optimization are achieved.
Patent Information
- Application Number
- CN202510432619.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-08-01
AI Technical Summary
The prior art is difficult to accurately identify the load state in a ball mill, resulting in low energy utilization, low efficiency, high energy consumption, and traditional methods are difficult to operate efficiently in embedded devices.
The multimodal signal fusion and attention enhancement method is used to identify the ball mill load through eigenmodal decomposition (FMD) and symmetric point generation mode (SDP) image generation technology, combined with the attention mechanism's lightweight network model (CBAM-EfficientNetV2).
It realizes high-precision ball mill load status recognition, improves ore grinding efficiency, reduces energy consumption, and is suitable for efficient operation in resource-constrained embedded equipment.
Smart Images

Figure CN120411701A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of monitoring and optimization of the ore grinding process, and particularly relates to a ball mill load identification method based on multi-modal signal fusion and attention enhancement. Background Art
[0002] The ball mill is a key grinding equipment in mining enterprises. The cylinder rotates to drive the steel balls and ore particles to crush the materials under the action of centrifugal force and gravity. However, due to the low energy utilization rate (only 10% - 20%), problems such as underload, empty grinding, and overfilling often occur during the operation of the ball mill, resulting in low efficiency, high energy consumption, and reduced product quality. Research shows that the accurate identification of the ball mill load state is of great significance for improving the ore grinding efficiency and reducing energy consumption.
[0003] However, due to the non-linearity and non-stationarity of the ball mill vibration signal, traditional methods (such as wavelet transform and empirical mode decomposition) are difficult to accurately extract signal features under complex working conditions and are easily affected by noise interference. And the existing load classification methods fail to fully utilize the multi-modal characteristics and spatial distribution characteristics in the signal, resulting in low classification accuracy. Moreover, many deep learning algorithms have high requirements for computing resources and are difficult to operate efficiently in embedded or industrial devices. Therefore, the present invention proposes a high-precision ball mill load identification method to accurately detect the load state of the ball mill by mining the characteristic information contained in the vibration signal of the ball mill cylinder, which helps to improve the ore grinding efficiency, achieve energy conservation and consumption reduction, and equipment intelligentization. Summary of the Invention
[0004] In order to overcome the deficiencies of the prior art, the present invention provides a ball mill load identification method based on multi-modal signal fusion and attention mechanism, which can extract effective signal features through the Feature Modal Decomposition (FMD) technology, combine the multi-modal feature fusion of the Symmetric Point Generation Pattern (SDP) image, and use the attention mechanism to enhance the key feature capture ability of the classification model, so as to realize the efficient identification of the ball mill load state, optimize the operating conditions, reduce energy consumption and improve production efficiency.
[0005] The specific steps of the invention are as follows: Step 1: Collect the vibration signal of the ball mill cylinder; Step 2: Preprocess the collected vibration signal; Step 3: Generate a multi-modal information fusion image as a feature image dataset; Step 4: Establish a ball mill load identification model based on attention enhancement; Step 5: Train the model and use the trained model to identify the ball mill load state for the SDP feature image dataset; Further, in step one, the acquisition source of the vibration signal data of the ball mill cylinder is from conducting experiments on a laboratory Bond work index dry ball mill, setting the acquisition frequency at 20KHz, and collecting the vibration signal of the ball mill cylinder through a DH5922N dynamic data acquisition instrument and a DH131 acceleration sensor.
[0006] Further, in step two, empirical mode decomposition (EMD) is used to perform denoising preprocessing on the original signal and extract effective signal features.
[0007] Further, the EMD method is a filtering-based decomposition method, the purpose of which is to decompose a complex signal into several modal components , each modal component corresponding to a different frequency range. This method realizes the separation of modal components through iterative filtering, and its main steps are as follows: Step A: Use a finite impulse response (FIR) filter to decompose the signal, each filter corresponding to a frequency range, and the design formula of the filter is:
[0008] where, is the Hanning window function, N is the window length, is the time point, is the frequency; Step B: The filter filters the original signal through convolution operation to extract the modal component of the k-th frequency range [[ID=3—1]]:
[0009] where, is the impulse response of the filter; Step C: Decompose the signal layer by layer through k filters, and reconstruct the original signal by superimposing the modal components to verify the accuracy and losslessness of the decomposition:
[0010] Step D: Retain the decomposed modal data finally obtained.
[0011] Further, in step three, the construction of the multi-modal information fusion image data set adopts the signal segmentation combined with symmetric point image (SDP) generation technology, and its main steps are as follows: Step A: Divide the long time series signal into segments of a fixed length for subsequent image generation or feature analysis. This method can uniformly extract the local characteristics of the signal and ensure the consistency of subsequent data processing; Step B: Scale the signal data in Step A to a unified range for signal normalization so that it can be represented with the same scale in the polar coordinate system: where, and represent the maximum and minimum values of the signal of channel k respectively; is the normalized signal, and its range is [0, 1]; Step C: Control the time offset of the signal data through the time delay τ to simulate the delay relationship between different time steps of the signal: where, represents the signal after being processed by the time delay τ; Step D: The signal data after the time delay processing is mapped into the polar coordinate system, and the mapping of the polar coordinates consists of two components, the radius and : where, is the base angle of each channel. Usually, the signal channels are evenly distributed within 360°; AF is the gain angle factor, which controls the change amplitude of the angle; Step E: Generate two symmetric polar coordinate points for the data of each signal channel: , , and the two angles reflect the forward and reverse changes of the signal, satisfying the following relationship: Step F: Map the signal of each channel into the polar coordinate system and use different angle values and normalized radii to plot on the same graph; each channel corresponds to a unique base angle , and the signal data of each channel forms multiple data points by calculating and . For the data of all channels, the final image consists of multiple sets of polar coordinate points: where, and represent the points of the signal in the polar coordinate system.
[0012] Furthermore, in Step Four, the ball mill load identification model adopts a lightweight network model based on attention enhancement, which can perform feature extraction and classification on the obtained data set.
[0013] Furthermore, the attention enhancement part in the lightweight network model based on attention enhancement adopts the CBAM method. The channel attention mechanism and spatial attention mechanism of the CBAM module are applied to the feature maps of each modal component, selectively highlighting key information through weighting and suppressing unimportant features, thereby improving the accuracy and robustness of the model.
[0014] Furthermore, the lightweight network part in the lightweight network model based on attention enhancement adopts the EfficientNetV2 network architecture. This network structure simultaneously optimizes the depth, width, and input resolution of the network through a compound scaling method, ensuring high accuracy at a low computational cost and being suitable for running on resource-constrained embedded devices.
[0015] Furthermore, in step five, the model training uses the constructed ball mill load identification model to classify the fused and enhanced feature images, uses the Softmax activation function to convert the output of the model into a probability distribution, and then judges the load state of the ball mill. The category with the highest probability is selected as the final prediction result, and the output categories include three states: underload, normal load, and overload.
[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: Comprehensive extraction and fusion of multi-modal features: The present invention combines feature modal decomposition (FMD) and symmetric point image (SDP) generation technology, and can effectively extract multi-modal features from the vibration signal and sound signal of the ball mill. Through FMD decomposition, high-order feature components in the signal are extracted, and further converted into image form through SDP image generation technology, thereby providing a richer and more comprehensive input for the subsequent classification model.
[0017] Attention mechanism improves classification accuracy: The present invention adopts the CBAM attention mechanism, which can adaptively focus on key regions in the input feature map, suppress interference from irrelevant regions, and thus improve the model's ability to capture load state features. The CBAM module focuses the model's attention on the most informative parts through two sub-modules, channel attention and spatial attention, significantly improving the classification accuracy and robustness.
[0018] Efficient and highly adaptable lightweight architecture: The present invention adopts EfficientNetV2 as the core network architecture. Compared with traditional convolutional neural networks (CNNs), EfficientNetV2 has higher computational efficiency and lower computational resource consumption. Through network structure optimization, EfficientNetV2 significantly reduces the amount of computation and storage requirements while ensuring the accuracy of the model, is suitable for running in embedded systems or hardware environments with limited resources, and has broad application prospects.
[0019] Excellent adaptability and robustness: The present invention organically combines a variety of advanced technologies (FMD, SDP, CBAM, EfficientNetV2), making full use of their respective advantages to effectively cope with the variability in the operating environment of the ball mill, including problems such as changes in load status, environmental noise interference, and signal instability. Through the combination of multi-modal feature extraction and attention mechanism, the model can maintain a high recognition accuracy under actual complex working conditions and has strong adaptability and robustness to environmental changes. Description of the Drawings
[0020] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are only one embodiment of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0021] Figure 1 It is a flowchart of a ball mill load recognition method based on multi-modal information fusion and attention enhancement in the embodiments of the present invention; Figure 2 It is a flowchart of the feature modal decomposition FMD in the embodiments of the present invention; Figure 3 It is an SDP diagram designed in the embodiments of the present invention; Figure 4 It is a network structure diagram of the mill load recognition model designed in the embodiments of the present invention; Figure 5 It is a CBAM attention mechanism structure diagram in the embodiments of the present invention; Figure 6 It is a confusion matrix diagram in the embodiments of the present invention. Detailed Embodiments
[0022] To make the technical means, creative features, achieved purposes and effects of the present invention easy to understand, the following clearly and completely describes the technical solutions in the specific embodiments of the present invention to further elaborate the present invention. Obviously, the described specific embodiments are only a part of the embodiments of the present invention, rather than all of them.
[0023] This specific embodiment is a ball mill load recognition method based on multi-modal information fusion and attention enhancement. The method flowchart is as Figure 1 shown, and the specific steps are as follows: Step 1: Conduct experiments using a laboratory Bond work index dry ball mill. The ore material used in the experiment is tungsten ore with a density of 1800 Kg / m3. Set the acquisition frequency to 20 KHz, and collect the vibration signals of the ball mill cylinder through a DH5922N dynamic data acquisition instrument and a DH131 acceleration sensor to obtain the original vibration signal data.
[0024] Step 2: Use empirical mode decomposition to decompose the original vibration signal into multiple modal components. The method flow chart is as Figure 2 shown. First, load the original signal and input two parameters: the number of modes n and the filter length L. Then, initialize the FIR filter bank through a Hanning window, use k filters, and start iteration. The design formula of the filter is as follows:
[0025] where, is the Hanning window function, N is the window length, is the time point, is the frequency; After that, the filter filters the original signal through convolution operation to obtain the modal component in the k-th frequency range:
[0026] Then, estimate the period of the mode using the original signal, and estimate the period of the mode by calculating the autocorrelation spectrum of the mode. The autocorrelation function is used to measure the similarity of the modal signal at different time delays:
[0027] Then, estimate the period autocorrelation spectrum of the mode and find the local maximum by finding the local maximum of the autocorrelation function to update the filter coefficients and complete one iteration:
[0028] After that, judge whether the number of iterations has reached the pre-set number of iterations. If not, return to step (3); if so, continue to input. Then, calculate the correlation matrix CC between every two modes to measure the similarity between the modes:
[0029] After calculating the correlations of all modes, select the mode with the highest correlation as the final retained mode (if the correlation of the mode is low, it will be discarded). In this way, continuously reduce the number of modes until the preset number of modes n is reached. Then, judge whether the mode k has reached the specified n. If not, return to step 3; otherwise, enter step 8; Finally, the retained mode is obtained as the final decomposition mode and output 、 、 、 。
[0030] Step 3: Process the decomposed modal components with the SDP algorithm to generate a multi-modal information fusion image as a feature image dataset; first, given the modal component signal and the segmentation length T (in the present invention, 1024 points are used as the signal length), the signal is divided into N non-overlapping segments:
[0031] When the signal length N cannot be divided evenly by T, the remaining part of the incomplete segment is padded with zeros:
[0032] After segmentation, the signal is converted into a two-dimensional matrix:
[0033] where K is the number of modes; N is the number of segments for each mode.
[0034] Scale the segmented signal data to a unified range for signal normalization so that it can be represented with the same scale in the polar coordinate system: where, and represent the maximum and minimum values of the signal of channel k respectively; is the normalized signal, and its range is [0, 1].
[0035] Then, control the time shift of the time delay τ of the signal data to simulate the delay relationship between signals at different time steps:
[0036] where, represents the signal after being processed by the time delay τ.
[0037] After the signal data is normalized and processed by time delay, it is mapped into the polar coordinate system. The mapping of the polar coordinate consists of the radius and two components, and its formula is as follows:
[0038] where, is the base angle of each channel, and the signal channels are usually evenly distributed within 360°; AF is the gain angle factor, which controls the variation range of the angle. Generate two symmetric polar coordinate points for the data of each signal channel: 、 , and the two angles reflect the forward and reverse changes of the signal, satisfying the following relationship:
[0039] Finally, map the signals of each channel into the polar coordinate system and use different angle values and normalized radii to plot on the same graph:
[0040] Among them, and represent the points of the signal in the polar coordinate system.
[0041] Finally, obtain the SDP map of multimodal information fusion, as shown in Figure 3 , and make all the obtained SDP maps into a dataset for subsequent processing.
[0042] Step 4. Divide the obtained multimodal information fusion image dataset into a training set, a validation set, and a test set according to the ratio of 7:2:1 for the classification training and effect verification of the subsequent mill load recognition model.
[0043] Step 5. Establish a lightweight network model CBAM-EfficientNetV2 based on attention enhancement. The network structure of the model is as shown in Figure 4 , and the main implementation process is as follows: 1. Input the training set divided from the SDP image dataset. These images contain the time-frequency features of the ball mill load state.
[0044] 2. Accept the input image , and extract the preliminary features through 3×3 convolution:
[0045] 3. Design the Fused-MBConv and MBConv modules for convolutional feature extraction and downsampling. The specific steps are as follows: (1) Expansion convolution: Increase the number of channels and enhance the expression ability of the model; (2) Depthwise separable convolution: Use depthwise convolution and pointwise convolution to reduce the computational amount; (3) Downsampling: Each MBConv module can perform downsampling through convolution with a stride of 2 to reduce the spatial dimension of the feature map; (4) Embedding the CBAM attention mechanism: After each MBConv module, embed CBAM (channel and spatial attention mechanism) to enhance the attention to important features. The structure of the CBAM attention mechanism is as shown in Figure 5 . First, perform global average pooling and global max pooling on the input features respectively:
[0046] where GAP is to calculate the average value of each channel; GMP is to calculate the maximum value of each channel; the output , .
[0047] Secondly, pass the pooling results through fully connected layers and non-linear activation respectively:
[0048] where and are the weight matrices of the fully connected layers; is the Sigmoid activation function, and the output .
[0049] Then, weight the original feature F through the channel weight :
[0050] Then, calculate the average value and maximum value of the channels of the input feature to generate the spatial attention input:
[0051] where Avg and Max calculate the average value and maximum value of C channels respectively; the output .
[0052] Use a 1×1 convolution to generate the spatial weight:
[0053] where represents the activation function Sigmoid, and the output ; Finally, weight the feature map through the channel weight to finally output the attention-weighted feature : 4. After passing through several MBConv modules, a feature map with a size of is obtained , and perform global average pooling on the output of the last MBConv module to compress the information in the spatial dimension into global information, providing input features for the classification layer:
[0054] This step reduces the feature map from to output ; 5. Flatten the features after global average pooling into a one-dimensional vector, and then perform classification through a fully connected layer and a Softmax activation function to output the load status of the ball mill (overload, normal load, underload)
[0055] where y is the classification result, representing the probability distribution of the load status; 6. By selecting the class with the highest probability in y, output the prediction result of the load status of the ball mill;
[0056] After establishing the mill load recognition model, set the network parameters of the pre-trained model as shown in Table 1 and train on the CBAM-EfficientNetV2 network model:
[0057] To verify the reliability of the ball mill load recognition method with multi-modal information fusion and attention enhancement of the present invention, use the test set to verify the classification effect of the constructed model, and draw the corresponding confusion matrix to evaluate the constructed model as Figure 6 shown. The confusion matrix is generated by the confusion_matrix(true_labels, pred_labels) function, where true_labels and pred_labels are the true labels and predicted labels of all test samples. This function returns an N×N matrix, and N is the number of classes. For the three-class classification problem (normal, overload, underload) in the present invention, the structure of the confusion matrix is shown in Table 2:
[0058] where TP is the number of samples correctly predicted by the model (the elements on the diagonal); FP is the number of samples mispredicted by the model as a certain class (the elements in the non-diagonal columns); FN is the number of samples that the model mispredicts a certain class as other classes (the elements in the non-diagonal rows); TN is the number of samples of other classes correctly predicted as other classes (the remaining part).
[0059] Then, the confusion matrix is normalized. The normalization is achieved by dividing each element in a row by the sum of that row, thereby obtaining the prediction percentage for that category:
[0060] Finally, the accuracy is calculated by computing the proportion of all correctly predicted samples to the total number of samples, and the performance of the model is evaluated based on the accuracy:
[0061] Step Six: After the model training is completed, save the trained mill load identification model; Step Seven: The trained model can be used to identify the load of the vibration signals of the ball mill collected under different working conditions.
[0062] The above describes the main technical features, basic principles, and related advantages of the present invention. For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary specific embodiments, and without departing from the concept or basic features of the present invention, the present invention can be implemented in other specific forms. Therefore, from any perspective, the above specific embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be encompassed within the present invention.
[0063] In addition, it should be understood that although this specification is described according to each embodiment, not every embodiment only contains an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A ball mill load identification method based on multi-modal signal fusion and attention enhancement, characterized in that The algorithm includes the following steps: Step 1: Collect the vibration signals of the ball mill cylinder; Step 2: Preprocess the collected vibration signals; Step 3: Generate a multi-modal information fusion image as a feature image dataset; Step 4: Establish a ball mill load recognition model with a fusion and attention mechanism; Step 5: Train the model, and use the trained model to recognize the ball mill load state of the feature image, and complete the monitoring and optimization of the ore grinding process.
2. The ball mill load identification method based on multi-modal signal fusion and attention enhancement according to claim 1, characterized in that In Step 1, the collection of the vibration signals of the ball mill cylinder is carried out by using a laboratory-type Bond work index dry ball mill for experiments.
3. The ball mill load identification method based on multi-modal signal fusion and attention enhancement according to claim 2, characterized in that, The ore material used in the experiment is tungsten ore with a density of 1800 Kg / m 3 ; The vibration signal of the mill cylinder is collected by a DH5922N dynamic data collector and a DH131 acceleration sensor. The collection frequency of the vibration signal is 20 KHz, and the signals of three mill load states, namely underload, normal load, and overload, are collected.
4. The ball mill load identification method based on multi-modal signal fusion and attention enhancement according to claim 1, characterized in that In Step 2, the signal preprocessing method is to use feature mode decomposition (FMD) to denoise the signal and extract effective signal features. This method realizes the separation of modal components through iterative filtering, and its main steps are as follows: Step A: Decompose the signal using a Finite Impulse Response (FIR) filter. Each filter corresponds to a frequency range. The design formula of the filter is: Among them, is the Hanning window function, N is the window length, is the time point, is the frequency; Step B: The filter filters the original signal through convolution operation to extract the modal component in the k-th frequency range : Among them, is the impulse response of the filter; Step C: Decompose the signal layer by layer through k filters, and reconstruct the original signal by superimposing the modal components to verify the accuracy and losslessness of the decomposition: Step D: Retain the decomposed modal data obtained finally.
5. The ball mill load identification method based on multi-modal signal fusion and attention enhancement according to claim 1, wherein In Step 3, the construction of the multi-modal information fusion image dataset adopts a signal segmentation combined with symmetric point image (SDP) generation technology, and its main steps are as follows: Step A: Divide the long-time series signal into segments with a fixed length for subsequent image generation or feature analysis. This method can uniformly extract the local characteristics of the signal and ensure the consistency of subsequent data processing; Step B: Scale the signal data in Step A to a unified range for signal normalization, so as to have the same scale for representation in the polar coordinate system: Among them, and respectively represent the maximum and minimum values of the signal of channel k; is the normalized signal, and its range is [0, 1]; Step C: Control the time shift of the signal data through the time delay τ to simulate the delay relationship between different time steps of the signal: Among them, represents the signal after being processed by the time delay τ; Step D: The signal data after delay processing is mapped into the polar coordinate system. The mapping of the polar coordinates consists of two components, the radius and : Among them, is the base angle of each channel. Usually, the signal channels are evenly distributed within 360°; AF is the gain angle factor, which controls the variation range of the angle; Step E: Generate two symmetric polar points for the data of each signal channel: , , where the two angles reflect the forward and reverse changes of the signal and satisfy the following relationship: Step F: Map the signals of each channel into the polar coordinate system and use different angular values and normalized radii to plot on the same graph; each channel corresponds to a unique base angle , and the signal data of each channel is calculated and to form multiple data points. For the data of all channels, the final image consists of multiple sets of polar coordinate points: Among them, and represent points of a signal in a polar coordinate system.
6. The ball mill load identification method based on multi-modal signal fusion and attention enhancement according to claim 1, wherein In Step 4, the ball mill load recognition model adopts a lightweight network model based on attention enhancement, which can extract features and classify the obtained dataset.
7. The ball mill load identification method based on multi-modal signal fusion and attention enhancement according to claim 6, characterized in that, The attention enhancement part in the lightweight network model based on attention enhancement adopts the CBAM method. The channel attention mechanism and spatial attention mechanism of the CBAM module are applied to the feature maps of each modal component, and key information is selectively highlighted by weighting, and unimportant features are suppressed, so as to improve the accuracy and robustness of the model.
8. The ball mill load identification method based on multi-modal signal fusion and attention enhancement according to claim 6, wherein The lightweight network part in the lightweight network model based on attention enhancement adopts the EfficientNetV2 network architecture. This network structure simultaneously optimizes the depth, width and input resolution of the network through a composite scaling method, ensuring high accuracy at a low computational cost and being suitable for running on resource-constrained embedded devices.
9. The ball mill load identification method based on multi-modal signal fusion and attention enhancement according to claim 1, wherein In Step 5, the model training uses the constructed ball mill load recognition model to classify the fused and enhanced feature images, uses the Softmax activation function to convert the output of the model into a probability distribution, and then judges the load state of the ball mill, and selects the category with the maximum probability as the final prediction result. The output categories include three states: underload, normal load and overload.
Citation Information
Cited By
Multi-target cooperative control method and system for powder grinding
CN121613859A
Ball mill internal state soft measurement method based on DEM simulation and sound signals
CN122364788A