Graphite ore image blurring elimination method based on deep learning
Through the GAN-based deblurring model and the Transformer module with multi-scale feature fusion, combined with the ECA attention mechanism, the motion blur problem in graphite ore image acquisition is solved, efficient and accurate ore grade identification is achieved, and equipment cost and computational complexity are reduced.
Patent Information
- Application Number
- CN202510651841.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-05-20
AI Technical Summary
Existing technologies suffer from poor image quality due to motion blur during graphite ore image acquisition, affecting subsequent segmentation and recognition work. In addition, high-pixel shooting equipment is expensive and difficult for companies to equip on a large scale. Existing deep learning methods perform poorly in practical applications.
A deblurring model based on a generative adversarial network (GAN) is adopted, combined with a Transformer module for transfer learning and multi-scale feature fusion, and the ECA attention mechanism is introduced. 7×7 convolutions are replaced by multiple 3×3 convolutions, bilinear interpolation upsampling is used, and data enhancement and the Adam optimizer are combined to gradually optimize the model to avoid overfitting.
It effectively eliminates motion blur in graphite ore images, improves image quality, enhances the accuracy and efficiency of ore grade identification, and reduces the overfitting risk and computational complexity of the model.
Smart Images

Figure CN120612253A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image recognition, and in particular relates to a method for deblurring graphite ore images based on deep learning. Background Art
[0002] During the mining and processing of graphite ore, effective differentiation of the grade of newly mined graphite ore is necessary to improve the efficiency of blending and beneficiation. Traditionally, there are two main methods for detecting and identifying the carbon grade of graphite ore. The first is traditional manual analysis, where workers rely on visual observation and experience to determine the approximate grade of the ore. This method is susceptible to subjective factors, resulting in low detection accuracy. Furthermore, long-term work in this manner can adversely affect worker health. The second method uses high-frequency combustion-infrared absorption spectroscopy to determine the fixed carbon content in ore samples. While this method offers high accuracy, it requires significant investment in manpower, time, and equipment wear and tear. With the rapid development of computer vision technology in recent years, it can rapidly process large amounts of ore image data, enabling automated analysis and identification of ore grades.
[0003] When using computer vision to analyze graphite ore grade, high-pixel cameras are first used to capture images of the ore moving on a belt. However, due to the complex lighting conditions and dusty environment of a production workshop, the captured images are often of poor quality and blurry. Furthermore, due to the high cost of high-pixel cameras, companies cannot deploy high-precision cameras on a large scale. The shooting environment is confined to the production workshop, where lighting is dim and dusty. The relative motion between the camera and the ore causes motion blur, which hinders subsequent segmentation and recognition. Therefore, finding an algorithm that can effectively process motion-blurred ore images is crucial for subsequent research on accurately identifying graphite ore grades and achieving efficient graphite ore grade classification. Summary of the Invention
[0004] According to a first aspect of the present invention, the present invention claims a method for deblurring a graphite ore image based on deep learning, comprising:
[0005] S1, constructing and generating a first adversarial network model, and pre-training the first adversarial network model on a sample data set;
[0006] S2, obtaining current model parameters of the first adversarial network model after pre-training, and saving the current model parameters to a binary file;
[0007] S3, constructing a first graphite ore image dataset, and dividing the first graphite ore image dataset into a training set and a validation set;
[0008] S4, inputting the training set into the first adversarial network model using the current model parameters for training and fine-tuning to obtain an improved second adversarial network model;
[0009] S5, using the second graphite ore image dataset to train the second adversarial network model, initializing model parameters, and obtaining a third adversarial network model;
[0010] S6. Evaluate the third adversarial network model based on the performance indicators of the validation set, and obtain a trained graphite ore blurred image removal model based on the evaluation results.
[0011] Furthermore, the S1 further includes:
[0012] Using a blurred graphite ore image as an input image of an input layer, multiple 3×3 convolutional layers are used to extract initial features of the input image;
[0013] The discriminator integrates an ECA module to perform channel attention weighting on the initial features of the input image and uses a dense residual block for feature learning and enhancement;
[0014] In the feature extraction and fusion stage, the Transformer module is introduced to enhance the feature extraction and fusion capabilities and handle long-distance dependencies;
[0015] The Transformer module captures global features through the self-attention mechanism, and then applies the ECA module again to enhance the representation ability of features;
[0016] The generator is used to output a clear image after deblurring, and pre-training is performed on the sample dataset; the pre-trained model parameters are saved in a binary file.
[0017] Furthermore, the S1 further includes:
[0018] In the generator, after each feature extraction and enhancement using a convolutional layer, an ECA module, a residual block, or a Transformer module, the feature map is upsampled using bilinear interpolation to gradually increase the resolution;
[0019] After multiple upsampling steps, the resolution of the feature map is the same as the input image, and the generator outputs a deblurred, clear image.
[0020] Adding activation layers and pooling layers in the generator and discriminator enhances the expressiveness of the model and reduces the amount of computation.
[0021] Furthermore, when the discriminator integrates the ECA module to perform channel attention weighting on the initial features of the input image and uses dense residual blocks for feature learning and enhancement, a weight is assigned to each output channel of each convolutional layer. The weight is obtained through adaptive learning and represents the degree of influence of each channel on the final output. The discriminator also includes:
[0022] The ECA module performs global average pooling on the initial features of the input image to obtain the average value of each channel;
[0023] The attention weight of each channel is obtained based on the average value of each channel, and the pooled features are modeled and normalized using one-dimensional convolution;
[0024] The channel attention weight is applied to the original input graphite ore image to complete the weighted operation.
[0025] Furthermore, before the step S4 is performed, the step further includes:
[0026] Set the loss function, initial learning rate, number of sample training batches, maximum number of training iterations, learning rate adjustment strategy, and gradient descent-based parameter optimization strategy;
[0027] Select Adam optimizer as the parameter optimization strategy based on gradient descent;
[0028] In the Adam optimizer, set the parameters β1 = 0.9, β2 = 0.99, where β1 is the weight decay strategy of the first-order moment estimate. The weight decay rate is used to limit the size of the model weight by adding a penalty term proportional to the square of the weight to the loss function.
[0029] Furthermore, the S4 further includes:
[0030] The graphite ore image dataset is divided into training set and validation set;
[0031] The pre-trained first adversarial network model is trained using the training set input, and a data augmentation strategy is used to improve the generalization ability of the model;
[0032] Fine-tune the model through adversarial training, including adversarial learning of the generator and discriminator;
[0033] After fine-tuning is completed, the improved second adversarial network model is obtained.
[0034] Furthermore, the S4 further includes:
[0035] The data enhancement strategies used include flipping, cropping, rotating, adjusting brightness and contrast, scaling, and erasing images;
[0036] Based on the discriminator's output probability for real images or the output probability for generated images being close to 0.5, the training of the generator and the discriminator is balanced, and the generator parameters are updated complementarily;
[0037] The relative blur loss function in the discriminator is expressed as:
[0038]
[0039] in, represents the relative blur loss function in the discriminator, and by minimizing the relative blur loss function, the discriminator learns to distinguish between real and generated images;
[0040] x represents a real image, which is obtained from the real data distribution p data (x) is sampled, represents the expected average value of all real images x, D(x) represents the output of the discriminator for the real image, and represents the probability that the discriminator believes that the real image is real, (D(x)-0.5) 2 Represents the squared error, which measures the difference between the discriminator's output probability of the real image and the target value 0.5;
[0041] z represents the noise input, from the noise distribution p z (z), G(z) represents the image generated by the generator based on the noise z, D(G(z)) represents the output of the discriminator for the generated image G(z), which is a probability value, indicating the probability that the discriminator believes that the image G(z) is real. represents the expected average value of all generated images G(z); (D(G(z))-0.5) 2 represents the squared error, which measures the difference between the discriminator's output probability of the generated image and the target value of 0.5.
[0042] Furthermore, the S5 further includes:
[0043] Using a graphite ore image dataset different from the graphite ore image dataset for training;
[0044] Initialize some model parameters and restart the training process;
[0045] The performance of the generator and discriminator is further optimized through adversarial training to obtain the third adversarial network model.
[0046] Furthermore, the S6 further includes:
[0047] During the model training process, the validation set is used for real-time performance evaluation;
[0048] Whenever the adversarial network model completes a training cycle, the performance of the adversarial network model is immediately tested using the validation set;
[0049] The training status of the model can be determined by observing whether the accuracy of the adversarial network model on the validation set continues to improve and whether the loss function continues to decrease.
[0050] If the accuracy continues to increase or the loss function continues to decrease, continue training; otherwise, if these indicators do not reach the expected trend, stop training.
[0051] The innovation of this invention compared with the existing invention methods is:
[0052] (1) Using deep learning methods, the blurred images of graphite ore are eliminated, thereby improving the rapid and accurate recognition of graphite ore grade; combining the transfer learning theory to train the model, the trained pre-parameters are efficiently transferred and reused, which can not only reduce the overfitting risk and training complexity of the model, but also improve the processing performance of the model;
[0053] (2) The original 7×7 convolution is replaced with multiple 3×3 convolutions in the model to reduce the number of parameters and increase the nonlinearity of the network. This improvement can improve the flexibility and learning ability of the network and reduce the amount of computation. At the same time, bilinear interpolation is used instead of transposed convolution for upsampling to avoid the checkerboard effect. Bilinear interpolation is a smoother upsampling method that can reduce artifacts in the generated image.
[0054] (3) A Transformer module based on multi-scale feature fusion is introduced to enhance feature extraction and fusion capabilities. The Transformer module can better handle long-range dependencies and improve feature representation capabilities. At the same time, a dense residual block (RRDB) is used to replace the traditional residual unit. RRDB fully utilizes the feature information of each layer through dense connections, enhancing the network's ability to learn image details.
[0055] (4) The traditional channel attention mechanism is replaced by the ECA channel attention mechanism. The ECA module is integrated into the residual block of the generator and applied at the end of each residual block to strengthen the channel attention. The ECA module can more efficiently utilize channel information and improve the expressiveness of features.
[0056] (5) By using a series of data enhancement methods such as random cropping, flipping, rotation, distortion, brightness and contrast adjustment, scaling, and erasing of images, combined with model training methods such as adversarial loss function and weight attenuation, the model fitting process can be accurately guided and the generalization performance of the model can be significantly enhanced. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 This is a schematic diagram showing the effect of eliminating motion blur in graphite ore in the prior art;
[0058] Figure 2 A flowchart of a method for eliminating motion blur in graphite ore based on deep learning as claimed in an embodiment of the present invention;
[0059] Figure 3 A structural diagram of a DeblurGAN-ore model for a method for eliminating motion blur in graphite ore based on deep learning as claimed in an embodiment of the present invention;
[0060] Figure 4 A schematic diagram of the ECA attention mechanism for a method for eliminating motion blur in graphite ore based on deep learning as claimed in an embodiment of the present invention;
[0061] Figure 5 A diagram showing the working principle of a GAN network for a method for eliminating motion blur in graphite ore based on deep learning, as claimed in an embodiment of the present invention. DETAILED DESCRIPTION
[0062] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0063] The terms "first", "second" and "third" in this application are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, a feature defined as "first", "second" and "third" may explicitly or implicitly include at least one of such features. In the description of this application, "multiple" means at least two, for example, two, three, etc., unless otherwise clearly and specifically defined. All directional indications in the embodiments of this application (such as up, down, left, right, front, back...) are only used to explain the relative positional relationship, movement, etc. between the components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally also include steps or units that are not listed, or may optionally also include other steps or units inherent to these processes, methods, products or devices.
[0064] References to "embodiments" herein mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0065] Traditional methods for detecting and identifying the carbon grade of graphite ore require a large investment of manpower, time, and equipment loss. In addition, the detection process is complex, with poor repeatability and a long detection time, usually more than several hours, which makes it difficult to meet the real-time sorting needs on the ore handling line.
[0066] The other category primarily relies on image analysis methods, though relatively few relevant results have been achieved. With the rapid development of computer vision technology in recent years, computer vision can rapidly process large amounts of ore image data, enabling automated analysis and identification of ore grades. In the process of using computer vision to analyze graphite ore grade, a large number of surface morphological images of the graphite ore are first collected and each image is labeled with the ore's corresponding grade, forming a relatively large dataset of labeled images. Artificial intelligence algorithms are then trained on this image dataset. The trained model can then automatically identify the specific grade of other unknown graphite ores. More specifically, when collecting graphite ore image datasets, the graphite ore is captured completely from random perspectives. The ore is then sent to a laboratory for chemical analysis to determine its carbon content and obtain an accurate grade classification, thus achieving a one-to-one correspondence between image and label. However, when using high-resolution cameras to photograph and sample the ore moving on a conveyor belt, the captured images are often of poor quality and blurry due to the complex lighting conditions and harsh dusty environment in the production workshop. Moreover, high-pixel cameras are expensive and limited by cost. The shooting environment is in a closed production workshop with dim light and high dust. The relative movement between the shooting equipment and the ore causes motion blur, which creates obstacles for subsequent segmentation and recognition work.
[0067] In addition, the cost of high-pixel shooting equipment is relatively high, and companies find it difficult to equip such equipment on a large scale for cost control reasons. In a closed production workshop, there is relative motion between the shooting equipment and the ore, and this motion causes motion blur in the image. Motion blur not only reduces the clarity of the image, but also brings additional difficulties to subsequent image segmentation and recognition, seriously affecting the accuracy and efficiency of ore identification, and thus restricting the intelligence level and automation level of the entire production process. Although the existing deep learning-based deblurring methods have made significant progress in image restoration, they also have some limitations and shortcomings, such as dependence on datasets, public datasets that do not match actual scenes, especially graphite ore datasets are more scarce. Most deep learning-based deblurring methods require pairs of blurred and clear images for training, but these synthetic datasets often have gaps with real-world blurred images, resulting in poor performance of the model in practical applications. Figure 1 The following diagram illustrates the effects of clear, blurred, and deblurred images of some graphite ores in existing technology. When images of graphite ore samples are blurred, optimizing them technically to improve image quality and, indirectly, enhance the accuracy of AI algorithms for identifying the grade of any graphite ore is a pressing technical challenge.
[0068] To address the aforementioned shortcomings or improvements in the existing technology, the present invention provides a deep learning-based method for deblurring graphite ore motion. Before using an artificial intelligence algorithm to learn on an image dataset, motion blur is removed from collected graphite ore images to improve data quality and subsequent grade identification accuracy. This specially designed model consists of a Transformer based on multi-scale feature fusion, a ConvNeXt Block, a global mean pooling layer, a normalized linear layer, a residual network, and a SoftMax layer. It operates by introducing a deblurring model based on a generative adversarial network (GAN). The model parameters are initialized using pre-trained weights from DeblurGAN-ore through transfer learning. Furthermore, to further deblur the image and highlight ore image details, the generator incorporates an ECA attention mechanism to increase the model's focus on key features and suppress irrelevant features, thereby improving network performance. The ECA module is optionally integrated into the generator or discriminator. The ECA module is added after each convolutional block to enhance feature representation. During GAN training, the ECA module is automatically included in the forward propagation of the generator and discriminator, requiring no additional effort.
[0069] According to the first embodiment of the present invention, the present invention claims a method for eliminating blur of graphite ore images based on deep learning, referring to Figure 2 ,include:
[0070] S1, constructing and generating a first adversarial network model, and pre-training the first adversarial network model on a sample data set;
[0071] S2, obtaining current model parameters of the first adversarial network model after pre-training, and saving the current model parameters to a binary file;
[0072] S3, constructing a first graphite ore image dataset, and dividing the first graphite ore image dataset into a training set and a validation set;
[0073] S4, inputting the training set into the first adversarial network model using the current model parameters for training and fine-tuning to obtain an improved second adversarial network model;
[0074] S5, using the second graphite ore image dataset to train the second adversarial network model, initializing model parameters, and obtaining a third adversarial network model;
[0075] S6. Evaluate the third adversarial network model based on the performance indicators of the validation set, and obtain a trained graphite ore blurred image removal model based on the evaluation results.
[0076] The pre-trained parameters of S1 form the basis for fine-tuning in S4, which in turn forms the basis for further training in S5. This gradual optimization approach allows the model to gradually transition from general capabilities to task-specific optimization.
[0077] S1 is pre-trained using a public general dataset, while S4 and S5 use the graphite ore image dataset, but S5 may use different data batches or a further expanded dataset.
[0078] S1 is preliminary training, S4 is fine-tuning for specific tasks, and S5 is further optimization based on fine-tuning. Each training is closer to the final application scenario.
[0079] Through this phased training approach, the model can gradually improve its ability to deblur graphite ore images while avoiding the overfitting problem that may result from direct training on a specific dataset.
[0080] Among them, in this embodiment, for the graphite ore image deblurring task, a corresponding generative adversarial network is preliminarily built;
[0081] Generative Adversarial Networks (GANs) are a highly effective choice for image deblurring, primarily due to their unique generative capabilities and adversarial training mechanism. They can generate high-quality, sharp images that are visually closer to real images. To improve deblurring performance, the improved DeblurGAN-ore network replaces the 7×7 convolutions in the generator with multiple 3×3 convolutions, reducing the number of parameters and increasing the network's nonlinearity. Furthermore, bilinear interpolation is used instead of transposed convolutions for upsampling, avoiding checkerboard artifacts. Therefore, the highly effective GAN network is selected as the basic framework of the network, which primarily consists of a generator and a discriminator. The generator uses a feature pyramid network (FPN) to extract multi-scale features. A Transformer model based on multi-scale feature fusion is integrated into the basic GAN network. The Transformer module can be introduced into the generator to enhance feature extraction and fusion capabilities. Secondly, a dense residual block (RRDB) is used to replace the traditional residual unit. RRDB utilizes dense connections to fully utilize feature information from each layer, enhancing the network's ability to learn image details. Residual blocks are used for feature learning and enhancement, and the ECA module at the end of each residual block further strengthens channel attention. Finally, upsampling and convolution layers are used to restore the original image size to generate a clear image.
[0082] Furthermore, the S1 further includes:
[0083] Using a blurred graphite ore image as an input image of an input layer, multiple 3×3 convolutional layers are used to extract initial features of the input image;
[0084] The discriminator integrates an ECA module to perform channel attention weighting on the initial features of the input image and uses a dense residual block for feature learning and enhancement;
[0085] In the feature extraction and fusion stage, the Transformer module is introduced to enhance the feature extraction and fusion capabilities and handle long-distance dependencies;
[0086] The Transformer module captures global features through the self-attention mechanism, and then applies the ECA module again to enhance the representation ability of features;
[0087] The generator is used to output a clear image after deblurring, and pre-training is performed on the sample dataset; the pre-trained model parameters are saved in a binary file.
[0088] Furthermore, the S1 further includes:
[0089] In the generator, after each feature extraction and enhancement using a convolutional layer, an ECA module, a residual block, or a Transformer module, the feature map is upsampled using bilinear interpolation to gradually increase the resolution;
[0090] After multiple upsampling steps, the resolution of the feature map is the same as the input image, and the generator outputs a deblurred, clear image.
[0091] Adding activation layers and pooling layers in the generator and discriminator enhances the expressiveness of the model and reduces the amount of computation.
[0092] In this embodiment, the 7×7 convolution in the generator is replaced with multiple 3×3 convolutions to reduce the number of parameters, increase the nonlinearity of the network, improve the flexibility and learning ability of the network, and reduce the amount of calculation.
[0093] Bilinear interpolation is used instead of transposed convolution for upsampling to avoid the checkerboard effect, which causes artifacts in the generated image and affects image quality. Bilinear interpolation is a smoother upsampling method that can reduce artifacts in the generated image;
[0094] The Residual-in-Residual Dense Block (RRDB) replaces the traditional residual unit. RRDB fully utilizes the feature information of each layer through dense connections, enhancing the network's ability to learn image details, thereby generating higher quality images.
[0095] The improved DeblurGAN-ore model is trained on a public image dataset. Here, the pre-trained weights on the Gopro dataset are mainly used, and the model parameters are initialized by introducing the model pre-training weights through transfer learning. Since large-scale datasets of minerals are relatively rare, the dataset used for model training this time is independently collected graphite ore images. There are 3152 pairs of training images, 6304 images. Each pair has a clear image and a corresponding blurred image. The larger the scale of this dataset, the better the pre-training effect of the model. Of course, after the model is successfully pre-trained on a large dataset, it already has good image deblurring capabilities; refer to Figure 3 , is the DeblurGAN-ore model structure diagram;
[0096] Save the parameter values of the pre-trained DeblurGAN-ore model in a file for subsequent experiments on the graphite ore image dataset;
[0097] We continued to collect a batch of graphite ore images. Using high-resolution cameras, we captured real-time graphite ore images on a conveyor belt traveling at 1.4 m / s. We used artificial synthesis to simulate the motion blur of the ore and generate blurred images corresponding to clear images to form image pairs for training. We included images of various grades, angles, and lighting conditions to expand the dataset and ensure it was large enough and had broad coverage.
[0098] The graphite mine image dataset is divided into a training set and a validation set. The training set is used to train the model's image deblurring capability and improve the model's generalization ability; the training set is used to verify the model's performance and determine whether to continue improving the model or terminate training.
[0099] To enhance the real-time image deblurring performance of the model, it is used for real-time image processing on ore conveyor belts. The traditional channel attention mechanism is replaced by the ECA channel attention mechanism. The traditional channel attention mechanism usually relies on fully connected layers or complex calculations to learn channel weights. These methods are often computationally intensive and have high model complexity. In order to solve the trade-off between performance and complexity, ECA proposes a one-dimensional convolution to replace the traditional fully connected layer to improve computational efficiency and parameter efficiency. Figure 5 ECA (Efficient Channel Attention) is an efficient channel attention mechanism used to enhance the feature representation capabilities of neural networks in the channel dimension. ECA adaptively weights the feature maps of each channel, allowing the network to pay more attention to the features of important channels, thereby improving network performance. The key point of ECA is to improve the expressiveness of features by introducing a lightweight attention mechanism to model the relationship between channels.
[0100] Furthermore, when the discriminator integrates the ECA module to perform channel attention weighting on the initial features of the input image and uses dense residual blocks for feature learning and enhancement, a weight is assigned to each output channel of each convolutional layer. The weight is obtained through adaptive learning and represents the degree of influence of each channel on the final output. The discriminator also includes:
[0101] The ECA module performs global average pooling on the initial features of the input image to obtain the average value of each channel;
[0102] The attention weight of each channel is obtained based on the average value of each channel, and the pooled features are modeled and normalized using one-dimensional convolution;
[0103] The channel attention weight is applied to the original input graphite ore image to complete the weighted operation.
[0104] In this embodiment, global average pooling averages all pixel values of each channel to obtain a scalar value for each channel as the feature representation of the channel. Assuming that the input feature map size is H×WH\timesWH×W (height and width), for each channel, GAP reduces its dimension to a scalar:
[0105]
[0106] Among them, x ijc Represents the pixel value of the input feature map at position (i, j), and y c is the average value of the cth channel;
[0107] The input of the one-dimensional convolution is the average value of each channel obtained by global average pooling (i.e. y c ), the output is the attention weight of each channel. The size of the convolution kernel is usually a hyperparameter, and its size does not change according to the size of the input image. Use one-dimensional convolution to model the pooled features:
[0108]
[0109] The output obtained by convolution needs to be further normalized, usually using a Sigmoid activation function to map the output value to the range [0, 1]. This is to ensure that the obtained weights can be used as weighting coefficients between channels, allowing the model to adjust the contribution of each channel according to these weights.
[0110]
[0111] Among them, σ(·) represents the Sigmoid activation function, α c is the attention weight of the c-th channel.
[0112] For each channel, the corresponding weight is applied to adjust the feature map of that channel. Adding the ECA module to each convolutional block in this network effectively enhances the network's feature representation capabilities. By adaptively weighting the feature maps of different channels, the model focuses on key features, improving its ability to learn and capture important information. ECA's efficient design reduces computational overhead through one-dimensional convolution operations, improving computational efficiency and making it particularly suitable for applications in resource-constrained environments. Furthermore, ECA enhances the network's robustness and generalization, helping the model maintain greater stability and accuracy in the face of noise and data variations, thereby improving overall network performance. In the generator of a GAN network, the ECA module can be used to enhance the feature representation of intermediate layers, resulting in generated images with richer texture and detail. By focusing on more important channels, the generator can more effectively utilize its capacity to produce higher-quality images. Here, we consider integrating the ECA module into the residual block of the generator. At the end of each residual block, the ECA module is applied to strengthen channel attention, and the result is then added to the block's input. In this way, the feature map obtained after convolution will be adjusted according to the importance of each channel, allowing the model to pay more attention to features that are useful for the task and suppress irrelevant information.
[0113] Furthermore, before the step S4 is performed, the step further includes:
[0114] Set the loss function, initial learning rate, number of sample training batches, maximum number of training iterations, learning rate adjustment strategy, and gradient descent-based parameter optimization strategy;
[0115] Select Adam optimizer as the parameter optimization strategy based on gradient descent;
[0116] In the Adam optimizer, set the parameters β1 = 0.9, β2 = 0.99, where β1 is the weight decay strategy of the first-order moment estimate. The weight decay rate is used to limit the size of the model weight by adding a penalty term proportional to the square of the weight to the loss function.
[0117] In this embodiment, the GAN model is trained on the graphite mine image dataset. Before starting the training model, the loss function is set and the initial learning rate is set to 1×10 -3This learning rate enables rapid updates of model parameters in the early stages of training, accelerating model convergence. As training progresses, the learning rate may be dynamically adjusted based on the learning rate adjustment strategy to ensure more precise parameter optimization in the later stages and avoid prematurely falling into local optima. The Adam optimizer is selected as the gradient descent-based parameter optimization strategy for the sample training batch size, maximum number of training iterations, learning rate adjustment strategy, and gradient descent-based parameter optimization strategy. The Adam optimizer combines the advantages of momentum and adaptive learning rate adjustment to effectively accelerate training and improve model convergence. In the Adam optimizer, parameters β1 = 0.9 and β2 = 0.99 are set. β1 is the decay rate of the first-order moment estimate, used to calculate the moving average of the gradient; β2 is the decay rate of the second-order moment estimate, used to calculate the moving average of the squared gradient. These parameter settings enable the Adam optimizer to adaptively adjust the learning rate during training while maintaining the stability of gradient updates. A weight decay rate w = 0.9 is used as the weight decay strategy. Weight decay is a regularization technique that limits the size of model weights by adding a penalty term proportional to the square of the weight to the loss function. This helps prevent overfitting and improves the model's generalization ability.
[0118] Furthermore, the S4 further includes:
[0119] Said S4 further comprises:
[0120] The graphite ore image dataset is divided into training set and validation set;
[0121] The pre-trained first adversarial network model is trained using the training set input, and a data augmentation strategy is used to improve the generalization ability of the model;
[0122] Fine-tune the model through adversarial training, including adversarial learning of the generator and discriminator;
[0123] After fine-tuning is completed, the improved second adversarial network model is obtained.
[0124] The data enhancement strategies used include flipping, cropping, rotating, adjusting brightness and contrast, scaling, and erasing images;
[0125] Based on the discriminator's output probability for real images or the output probability for generated images being close to 0.5, the training of the generator and the discriminator is balanced, and the generator parameters are updated complementarily;
[0126] The relative blur loss function in the discriminator is expressed as:
[0127]
[0128] in, represents the relative blur loss function in the discriminator, and by minimizing the relative blur loss function, the discriminator learns to distinguish between real and generated images;
[0129] x represents a real image, which is obtained from the real data distribution p data (x) is sampled, represents the expected average value of all real images x, D(x) represents the output of the discriminator for the real image, and represents the probability that the discriminator believes that the real image is real, (D(x)-0.5) 2 Represents the squared error, which measures the difference between the discriminator's output probability of the real image and the target value 0.5;
[0130] z represents the noise input, from the noise distribution p z (z), G(z) represents the image generated by the generator based on the noise z, D(G(z)) represents the output of the discriminator for the generated image G(z), which is a probability value, indicating the probability that the discriminator believes that the image G(z) is real. represents the expected average value of all generated images G(z); (D(G(z))-0.5) 2 Represents the squared error, which measures the difference between the output probability of the discriminator for the generated image and the target value 0.5; Figure 4 This is a diagram of the working principle of the GAN network.
[0131] The S5 further includes:
[0132] Using a graphite ore image dataset different from the graphite ore image dataset for training;
[0133] Initialize some model parameters and restart the training process;
[0134] The performance of the generator and discriminator is further optimized through adversarial training to obtain the third adversarial network model.
[0135] Furthermore, the S6 further includes:
[0136] During the model training process, the validation set is used for real-time performance evaluation;
[0137] Whenever the adversarial network model completes a training cycle, the performance of the adversarial network model is immediately tested using the validation set;
[0138] The training status of the model can be determined by observing whether the accuracy of the adversarial network model on the validation set continues to improve and whether the loss function continues to decrease.
[0139] If the accuracy continues to increase or the loss function continues to decrease, continue training; otherwise, if these indicators do not reach the expected trend, stop training.
[0140] Among them, in this embodiment, a GAN artificial intelligence model is obtained after training. By inputting a graphite ore image with poor image quality into the model, a clear image with improved image quality can be obtained, and real-time deblurring processing of the ore image can be achieved, which effectively solves the problem of blurring of the graphite ore image, improves the image quality and the accuracy of graphite ore grade identification, and provides technical support for the subsequent real-time identification of graphite ore grade.
[0141] The following is described with specific examples:
[0142] Taking the blurring caused by graphite ore on a conveyor belt moving at a speed of 1.4 m / s as an example, the main steps of the invention are as follows:
[0143] The basic GAN model was built using the Python 3.10 programming language and the PyTorch 2.0.1 deep learning framework. The backbone network consists of Inception-ResNet-v2 and MobileNet. The former provides the highest accuracy, while the latter significantly reduces model parameters and computational complexity while maintaining high deblurring quality.
[0144] The improved DeblurGAN-ore model was pre-trained on the GOPRO dataset, allowing the model's reference indicators on the validation set, peak signal-to-noise ratio (PSNR), to reach approximately 28, and the structural similarity index (SSIM), to reach approximately 0.8. During training, the image pixels were standardized, and the mean and standard deviation values used were calculated based on the pixel values of all image samples in the training set.
[0145] Save the parameters of the generator and discriminator of the pre-trained DeblurGAN-ore model to a binary file.
[0146] Collect graphite ore images under different lighting conditions, with varying grades and motion speeds, ensuring a sufficiently large dataset to form a graphite ore image dataset of sufficient size. Collect graphite ore images, including blurred images and corresponding sharp images (if available), to form a dataset of sufficient size. For each blurred graphite ore image, use appropriate preprocessing methods, such as removing as much background pixels as possible to improve model recognition efficiency and accuracy. Finally, scale the image to a minimum of 256 pixels on its shortest side.
[0147] The collected graphite ore image dataset is divided into training set and test set in a ratio of 7:3.
[0148] The saved parameters are loaded into the generator and discriminator of the DeblurGAN-ore model. Then, all samples of the graphite ore blurred image dataset are forward-operated in the DeblurGAN-ore model in turn, and the model is fine-tuned as needed.
[0149] The DeblurGAN-ore model is trained on the collected graphite ore image dataset. The loss function is the adversarial loss function and the relative blur loss function. The initial learning rate is set to 1×10 -3 The number of sample training batches was set to 32, the maximum number of training iterations was set to 200 training set cycles, and the learning rate adjustment strategy used an adaptive learning rate decay loss. This dynamically adjusts the learning rate based on performance indicators (such as the loss function value) during training to improve training efficiency and model performance. The Adam optimizer was used for parameter optimization, with β1 = 0.9, β2 = 0.99, and a weight decay rate w = 0.9. For data augmentation, the collected data images were flipped, cropped, rotated, and subjected to brightness and contrast adjustments to simulate sample data collected under different environments to enhance data diversity.
[0150] After each training cycle, the model uses the validation set to evaluate its accuracy and loss function trends. Training continues as long as the accuracy continues to improve or the loss function continues to decrease; conversely, if these indicators no longer improve, training stops.
[0151] After training, a model for removing blur from graphite ore images can be obtained. Its performance evaluation indicators, peak signal-to-noise ratio (PSNR), reach about 35, and structural similarity index (SSIM) reaches about 0.95.
[0152] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0153] In addition, the functional units in the various embodiments of the present application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units. The above is only an implementation method of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the description and drawings of this application, or directly or indirectly used in other related technical fields, are also included in the patent protection scope of the present application.
[0154] The above detailed description of the specific embodiments of the invention is intended only as an example, and the present application is not limited to the specific embodiments described above. For those skilled in the art, any equivalent modifications or substitutions of the invention are also within the scope of the present application. Therefore, equivalent changes, modifications, and improvements made without departing from the spirit and scope of the present application should be included within the scope of the present application.
[0155] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0156] In addition, the functional units in the various embodiments of the present application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units. The above is only an implementation method of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the description and drawings of this application, or directly or indirectly used in other related technical fields, are also included in the patent protection scope of the present application.
[0157] The above detailed description of the specific embodiments of the invention is intended only as an example, and the present application is not limited to the specific embodiments described above. For those skilled in the art, any equivalent modifications or substitutions of the invention are also within the scope of the present application. Therefore, equivalent changes, modifications, and improvements made without departing from the spirit and scope of the present application should be included within the scope of the present application.
Claims
1. A method for deblurring graphite ore images based on deep learning, characterized in that: include: S1, constructing and generating a first adversarial network model, and pre-training the first adversarial network model on a sample data set; S2, obtaining current model parameters of the first adversarial network model after pre-training, and saving the current model parameters to a binary file; S3, constructing a first graphite ore image dataset, and dividing the first graphite ore image dataset into a training set and a validation set; S4, inputting the training set into the first adversarial network model using the current model parameters for training and fine-tuning to obtain an improved second adversarial network model; S5, using the second graphite ore image dataset to train the second adversarial network model, initializing model parameters, and obtaining a third adversarial network model; S6. Evaluate the third adversarial network model based on the performance indicators of the validation set, and obtain a trained graphite ore blurred image removal model based on the evaluation results.
2. The method for deblurring a graphite ore image based on deep learning according to claim 1, characterized in that: Said S1 further includes: Using a blurred graphite ore image as an input image of an input layer, multiple 3×3 convolutional layers are used to extract initial features of the input image; The discriminator integrates an ECA module to perform channel attention weighting on the initial features of the input image and uses a dense residual block for feature learning and enhancement; In the feature extraction and fusion stage, the Transformer module is introduced to enhance the feature extraction and fusion capabilities and handle long-distance dependencies; The Transformer module captures global features through the self-attention mechanism, and then applies the ECA module again to enhance the representation ability of features; The generator is used to output a clear image after deblurring, and pre-training is performed on the sample dataset; the pre-trained model parameters are saved in a binary file.
3. The method for deblurring a graphite ore image based on deep learning according to claim 2, characterized in that: Said S1 further includes: In the generator, after each feature extraction and enhancement using a convolutional layer, an ECA module, a residual block, or a Transformer module, the feature map is upsampled using bilinear interpolation to gradually increase the resolution; After multiple upsampling steps, the resolution of the feature map is the same as the input image, and the generator outputs a deblurred, clear image. Adding activation layers and pooling layers in the generator and discriminator enhances the expressiveness of the model and reduces the amount of computation.
4. The method for deblurring a graphite ore image based on deep learning according to claim 2, characterized in that: When the discriminator integrates the ECA module to perform channel attention weighting on the initial features of the input image and uses dense residual blocks for feature learning and enhancement, a weight is assigned to each output channel of each convolutional layer. The weight is obtained through adaptive learning and represents the degree of influence of each channel on the final output. It also includes: The ECA module performs global average pooling on the initial features of the input image to obtain the average value of each channel; The attention weight of each channel is obtained based on the average value of each channel, and the pooled features are modeled and normalized using one-dimensional convolution; The channel attention weight is applied to the original input graphite ore image to complete the weighted operation.
5. The method for deblurring a graphite ore image based on deep learning according to claim 2, characterized in that: Before the S4 is performed, the following steps are further included: Set the loss function, initial learning rate, number of sample training batches, maximum number of training iterations, learning rate adjustment strategy, and gradient descent-based parameter optimization strategy; Select Adam optimizer as the parameter optimization strategy based on gradient descent; In the Adam optimizer, set the parameters β1 = 0.9, β2 = 0.99, where β1 is the weight decay strategy of the first-order moment estimate. The weight decay rate is used to limit the size of the model weight by adding a penalty term proportional to the square of the weight to the loss function.
6. The method for deblurring a graphite ore image based on deep learning according to claim 2, characterized in that: Said S4 further comprises: The graphite ore image dataset is divided into training set and validation set; The pre-trained first adversarial network model is trained using the training set input, and a data augmentation strategy is used to improve the generalization ability of the model; Fine-tune the model through adversarial training, including adversarial learning of the generator and discriminator; After fine-tuning is completed, the improved second adversarial network model is obtained.
7. The method for deblurring graphite ore images based on deep learning according to claim 2, characterized in that: Said S4 further comprises: The data enhancement strategies used include flipping, cropping, rotating, adjusting brightness and contrast, scaling, and erasing images; Based on the discriminator's output probability for real images or the output probability for generated images being close to 0.5, the training of the generator and the discriminator is balanced, and the generator parameters are updated complementarily; The relative blur loss function in the discriminator is expressed as: in, represents the relative blur loss function in the discriminator, and by minimizing the relative blur loss function, the discriminator learns to distinguish between real and generated images; x represents a real image, which is obtained from the real data distribution p data (x) is sampled, represents the expected average value of all real images x, D(x) represents the output of the discriminator for the real image, and represents the probability that the discriminator believes that the real image is real, (D(x)-0.5) 2 Represents the squared error, which measures the difference between the discriminator's output probability of the real image and the target value 0.5; z represents the noise input, from the noise distribution p z (z), G(z) represents the image generated by the generator based on the noise z, D(G(z)) represents the output of the discriminator for the generated image G(z), which is a probability value, indicating the probability that the discriminator believes that the image G(z) is real. represents the expected average value of all generated images G(z); (D(G(z))-0.5) 2 represents the squared error, which measures the difference between the discriminator's output probability of the generated image and the target value of 0.
5.
8. The method for deblurring graphite ore images based on deep learning according to claim 6, characterized in that: The S5 further includes: Using a graphite ore image dataset different from the graphite ore image dataset for training; Initialize some model parameters and restart the training process; The performance of the generator and discriminator is further optimized through adversarial training to obtain the third adversarial network model.
9. The method for deblurring a graphite ore image based on deep learning according to claim 2, characterized in that: The S6 further includes: During the model training process, the validation set is used for real-time performance evaluation; Whenever the adversarial network model completes a training cycle, the performance of the adversarial network model is immediately tested using the validation set; The training status of the model can be determined by observing whether the accuracy of the adversarial network model on the validation set continues to improve and whether the loss function continues to decrease. If the accuracy continues to increase or the loss function continues to decrease, continue training; otherwise, if these indicators do not reach the expected trend, stop training.
Citation Information
Patent Citations
Image watermark removing method based on adversarial network
CN111105336A
Depth image deblurring method based on multi-scale fusion coding network
CN113129237A
Mineral classification method based on deep convolution fusion multi-scale image features
CN116416479A
Method for generating infrared image based on dense residual error and attention guiding
CN118505833A
Method of analysing mineralogical composition of crystalline rocks
RU2834385C1