Ore size measurement and anomaly identification method and system

By designing the OreSizeNet network of a multi-task learning framework, combined with the depth separable convolution module (DWABlock), the problems of low accuracy, low efficiency and poor robustness in ore block detection are solved, and high-precision and real-time ore block measurement and abnormal recognition are achieved.

CN119516288BActive Publication Date: 2025-05-09CHANGSHA RES INST OF MINING & METALLURGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510089755.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-09
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

The prior art has problems such as low accuracy, low efficiency, poor robustness and difficulty in achieving multi-task learning in ore block detection, especially when facing complex backgrounds and abnormal detection, the detection accuracy and real-time performance are insufficient.

Method used

A deep convolutional neural network with a multi-task learning framework is designed, called OreSizeNet, which includes two types of task branches: binary classification tasks and multi-scale segmentation tasks. The network adopts standard convolution modules, feature fusion modules and depth separable convolution modules (DWABlocks), and realizes accurate measurement and abnormal identification of ore block size through multi-scale feature extraction and cross-layer feature fusion.

Benefits of technology

It improves the accuracy and efficiency of ore blocking detection, enhances the robustness of the model in complex environments, and can accurately identify ore blocking and abnormal changes under real-time conditions, which is significantly better than YOLOv5 and YOLOv8.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119516288B_ABST
    Figure CN119516288B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of neural network technology, and discloses a method and system for measuring and identifying ore block size, so as to improve the accuracy and efficiency of ore detection, and effectively handle complex situations in ore images. The method comprises: constructing, training and testing an OreSizeNet network, wherein the OreSizeNet network comprises two types of task branches, wherein the first type of task branch is a binary classification task based on whether the image frame as a whole is abnormal, and the second type of task branch is a segmentation task based on the pixel level of the image frame at three different scales of large, medium and small for calculating the block size of the ore, and the first and second types of task branches share part of the front network; wherein a DWABlock module is deployed in the part of the front network shared by the first and second types of task branches, in the intermediate network before each classification head after the first and second types of task branches are differentiated, and in each detection head.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of neural network technology, and in particular to a method and system for measuring ore size and identifying anomalies. Background Art

[0002] Ore size detection plays a vital role in mining production, especially in the process of grinding, crushing and screening. Accurate particle size monitoring is directly related to production efficiency, equipment service life and energy consumption. Although traditional ore size detection methods, such as manual screening and mechanized measurement, can provide particle size information to a certain extent, they have many shortcomings. First of all, manual screening is not only time-consuming and labor-intensive, but also easily interfered by human factors, resulting in inaccurate data. Secondly, traditional mechanized methods cannot obtain the changes in particle size in real time, and it is difficult to achieve fine classification of ore size, which brings potential production risks. Therefore, the development of a method that can detect ore size in real time and accurately has become an urgent problem to be solved in the mining field.

[0003] With the rapid development of computer vision and deep learning technology, ore size detection methods based on image recognition have gradually become a research hotspot. In particular, the widespread application of deep convolutional neural networks (CNNs) in image analysis has made ore size detection technology develop in the direction of automation and intelligence. Although existing deep learning models, such as U-Net, YOLOv5, and YOLOv8, have made significant progress in the fields of object detection and segmentation, their application in ore size detection still faces some challenges, especially in terms of anomaly detection of ore size, robustness in complex backgrounds, and computing resource consumption.

[0004] As a classic image segmentation network, U-Net has been widely used in the segmentation tasks of medical images and geological images. However, U-Net faces the problem of low computational efficiency in ore fragmentation detection. It usually requires a large amount of training data and high computing resources, especially when processing large-scale ore images, the time cost of training and inference is high. In addition, U-Net lacks the ability of multi-task learning and mainly focuses on pixel-level segmentation tasks. It is difficult to handle multi-task joint optimization problems such as ore fragmentation measurement and anomaly detection. Moreover, when faced with complex backgrounds such as noise and illumination changes in ore images, U-Net has poor adaptability and is prone to errors.

[0005] As an efficient target detection model, YOLOv5 is widely used in the field of object detection due to its strong real-time detection capability. However, the limitations of YOLOv5 in ore particle size detection are also obvious. It is suitable for the detection of larger objects, but it may not provide sufficient resolution for the detection of ore particles with fine particle size differences. When faced with ore accumulation, overlap or complex background interference, the detection accuracy and robustness of YOLOv5 may also be affected, especially in the detection of abnormal particle size. YOLOv5 is not optimized for abnormal changes, resulting in weak detection capabilities for abnormal particle size or abnormal fragmentation.

[0006] As the latest version of the YOLO series, YOLOv8 has made significant improvements in target detection accuracy and speed, especially in small object detection and refined recognition. However, YOLOv8 still faces some limitations in ore block size detection. Although it performs well in general object detection, it has poor adaptability to changes in ore block size, especially when the ore morphology is complex and varied, the detection effect of YOLOv8 may not meet the needs. In addition, although YOLOv8 has been optimized in speed and accuracy, its large model and complex reasoning process make it difficult to perform in real time and deploy in ore block size detection, especially in industrial environments with limited resources, the effect of real-time detection may also be affected. At the same time, YOLOv8 lacks optimization for ore block size anomaly detection, which may lead to missed detection or false detection.

[0007] In summary, the specific technical issues include:

[0008] 1. Accuracy and efficiency of ore size measurement: In the actual production process, the size, shape and other characteristics of the ore vary greatly, and traditional methods are difficult to achieve refined measurement. How to use an efficient deep learning network to measure the size of the ore with high accuracy, especially in the simultaneous identification of large and small pieces of ore, is still an urgent problem to be solved.

[0009] 2. Accuracy of identifying ore anomalies: During the sorting process of ore on the conveyor belt, some abnormal blocks may appear, such as cracks, impurity minerals, etc. Existing technologies often have difficulty in effectively distinguishing normal ores from abnormal ores, especially when the ore size and shape are complex and changeable, the accuracy and robustness of identifying abnormal minerals are insufficient.

[0010] 3. Challenges of multi-scale feature extraction: The target sizes involved in ore size measurement and anomaly identification vary greatly. How to design a network structure that can simultaneously process targets of different scales and improve the network's comprehensive recognition capabilities for ore size and anomalies is still a technical bottleneck for deep learning in this field.

[0011] 4. Efficient fusion of classification and segmentation tasks: In the tasks of ore size measurement and anomaly identification, how to effectively integrate classification and segmentation tasks so as to realize ore size classification and anomaly detection in the same network model while maintaining efficient computing performance is the key difficulty in realizing multi-task learning.

[0012] 5. Category imbalance and small target detection: In ore images, large pieces of ore usually occupy a larger area, while the proportion of small pieces of ore or abnormal ore is smaller. How to deal with this category imbalance problem and improve the detection accuracy of small targets is still a challenge that is difficult to overcome with existing technologies. Summary of the invention

[0013] The present invention aims to disclose a method and system for measuring ore size and identifying anomalies, so as to improve the accuracy and efficiency of ore detection and effectively handle complex situations in ore images.

[0014] To achieve the above object, the method of the present invention comprises:

[0015] Construct, train and verify the OreSizeNet network, which includes two types of task branches. The first type of task branch is a binary classification task based on whether the image frame is abnormal as a whole, and the second type of task branch is a segmentation task based on the pixel level of the image frame at three different scales: large, medium and small, for calculating the size of the ore. The first and second types of task branches share part of the front network.

[0016] Based on the OreSizeNet network, the real-time collected images are synchronously measured for ore size and anomaly identification;

[0017] The OreSizeNet network includes a standard convolution module for feature processing, a feature fusion module and a DWABlock module based on deep separable convolution; the DWABlock module is deployed in the front network shared by the first and second task branches; at least one DWABlock module is deployed in each intermediate network between the first and second task branches and before each classification head; and the DWABlock module is also deployed in each detection head;

[0018] The DWABlock module first performs initial feature transformation with a 1×1 standard convolution unit with a stride of 1, and then extracts local and global multi-scale features with three parallel paths and then fuses them. Each path uses two depth-wise separable convolutions to process the horizontal and vertical information in series. The first path is preceded and followed by two depth-wise separable convolutions of 1×3 and 5×1, respectively. The second path is preceded and followed by two depth-wise separable convolutions of 1×7 and 5×1, respectively. The third path is preceded and followed by two depth-wise separable convolutions of 5×1 and 1×7, respectively. The fused features are then output after the number of channels is adjusted by a 1×1 standard convolution.

[0019] Preferably, the classification heads of large, medium and small scales of the second task branch adopt SegHead modules, and the data processing flow of the SegHead module includes three stages: initial feature adjustment, multi-layer feature extraction and fusion, and final task prediction, specifically:

[0020] First, the input features are preliminarily adjusted through a 1×1 standard convolution, and then progressive feature extraction is performed using three combinations consisting of a DWABlock module and a feature fusion module. The first feature fusion module fuses the output of the first DWABlock module with the output of the preliminary adjustment standard convolution, the second feature fusion module fuses the output of the second DWABlock module with the output of the first DWABlock module, and the third feature fusion module fuses the output of the third DWABlock module with the output of the second DWABlock module. The output of the third feature fusion module is adjusted for the number of channels through a 1×1 standard convolution and then output for final classification prediction.

[0021] Preferably, the present invention further comprises:

[0022] The actual size of the ore is calculated based on the segmentation results of the OreSizeNet network in the second task branch, the internal and external parameters of the camera calibration, and the conversion relationship from pixel coordinates to actual coordinates.

[0023] Preferably, the classification head of the first type of task branch adopts a ClsHead module. In the ClsHead module, the input features are first adjusted to the number of channels through a 1×1 standard convolution, and the adjusted features are sent to the DWABlock module. The DWABlock output and the features of the standard convolution output with the adjusted channel number are fused and then adjusted to the number of channels through a 1×1 standard convolution for final classification prediction.

[0024] Preferably, the loss function of the OreSizeNet network manipulates the following four parts:

[0025] Part 1: Binary Cross Entropy Loss for Measuring the Difference Between Model Output and True Label ; The calculation formula is:

[0026] ;

[0027] in, Pixel The real label, is the predicted value, is the total number of pixels;

[0028] Part 2: Dice loss for solving category imbalance in segmentation tasks , the calculation formula is:

[0029] ;

[0030] in, is a smoothing factor used to prevent the denominator from being zero;

[0031] The third part, Tversky loss, is calculated as follows:

[0032] ;

[0033] in, and is a hyperparameter;

[0034] Part 4: Weighted Binary Cross Entropy Loss , the calculation formula is:

[0035] ;

[0036] Among them, the weight Dynamically adjust according to the area to which the pixel belongs;

[0037] The final comprehensive loss function is the weighted sum of the above losses, and its expression is:

[0038] ;

[0039] in, is the weight hyperparameter of each part of the loss.

[0040] To achieve the above objectives, the present invention also discloses a system for measuring ore size and identifying anomalies, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above method when executing the computer program.

[0041] The present invention has the following beneficial effects:

[0042] Compared with the traditional U-Net, YOLOv5 and YOLOv8, the ore block size detection method based on multi-task learning and deep convolutional neural network proposed in the present invention has significant advantages. The present method adopts a multi-task learning framework, which can simultaneously process the measurement and anomaly detection tasks of ore block size, thereby improving the comprehensiveness and accuracy of detection. Through joint optimization, the model can detect and identify abnormal changes in ore block size in real time while ensuring the accuracy of block size measurement. In addition, the present invention effectively improves the robustness of the model in complex environments through a unique DWABlock module, and can stably operate in complex scenes such as ore accumulation, illumination changes and background interference. Compared with the YOLO series models, the present invention is more accurate in identifying fine grains and abnormal changes in ore images. At the same time, in view of the characteristics of ore block size detection, the present invention is optimized in terms of computational efficiency, which is significantly better than YOLOv5 and YOLOv8. In the real-time detection of large-scale data sets, it can effectively reduce the consumption of computing resources and maintain a high detection accuracy. Moreover, it can accurately identify and classify abnormal changes in ore size based on the first type of task branch, avoiding problems such as uneven equipment load and low production efficiency.

[0043] Thus, the ore block size and anomaly detection method disclosed in the present invention not only overcomes the shortcomings of existing methods in terms of accuracy, robustness and real-time performance, but also can provide more accurate and efficient detection results in complex ore processing environments, and has broad application prospects.

[0044] The present invention will be further described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The drawings constituting a part of this application are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0046] Figure 1 It is a structural block diagram of the DWABlock module disclosed in the embodiment of the present invention.

[0047] Figure 2 It is an architecture block diagram of the OreSizeNet network disclosed in an embodiment of the present invention.

[0048] Figure 3 It is a structural block diagram of the SegHead module disclosed in an embodiment of the present invention.

[0049] Figure 4 It is a structural block diagram of the ClsHead module disclosed in the embodiment of the present invention. DETAILED DESCRIPTION

[0050] The embodiments of the present invention are described in detail below with reference to the accompanying drawings, but the present invention can be implemented in many different ways as defined and covered by the claims.

[0051] Embodiment 1:

[0052] The present embodiment discloses a method for measuring the size of ore and identifying anomalies, comprising: constructing, training and verifying an OreSizeNet network, wherein the OreSizeNet network comprises two types of task branches, wherein the first type of task branch is a binary classification task based on whether the image frame as a whole is abnormal, and the second type of task branch is a segmentation task based on the pixel level of the image frame at three different scales of large, medium and small for calculating the size of ore, and the first and second types of task branches share part of the front network; then, based on the OreSizeNet network, the ore size measurement and anomaly identification are synchronously performed on the real-time collected images.

[0053] In this embodiment, the OreSizeNet network includes a standard convolution module for feature processing, a feature fusion module, and a DWABlock module based on deep separable convolution. Figure 1 As shown, the DWABlock module of the present invention first performs initial feature transformation with a 1×1 standard convolution unit with a stride of 1, and then extracts local and global multi-scale features with three parallel paths and then fuses them. Each path uses two depth-separable convolutions to process the horizontal and vertical information in series. The first path is preceded and followed by two depth-separable convolutions of 1×3 and 5×1, respectively. The second path is preceded and followed by two depth-separable convolutions of 1×7 and 5×1, respectively. The third path is preceded and followed by two depth-separable convolutions of 5×1 and 1×7, respectively. The fused features are then output after the number of channels is adjusted by a 1×1 standard convolution.

[0054] In this embodiment, a DWABlock module is deployed in the front network shared by the first and second task branches; at least one DWABlock module is deployed in the intermediate network between the differentiation of the first and second task branches and before each classification head; and a DWABlock module is also deployed in each detection head.

[0055] In this embodiment, a specific OreSizeNet network structure is as follows: Figure 2 As shown, including:

[0056] 1. Feature extraction module process based on nodes 1-0 to 1-9.

[0057] The feature extraction module of OreSizeNet gradually extracts multi-scale features from the input image and realizes feature fusion through multiple downsampling and depthwise separable convolution modules (DWABlock). First, the input image is subjected to a 3×3 convolution with a stride of 1 (node ​​1-0) to extract the initial feature map F1, which enhances the feature expression capability while maintaining the spatial resolution. Subsequently, F1 undergoes another 3×3 convolution with a stride of 2 (node ​​1-1) to generate a downsampled feature map F2, which significantly reduces the spatial resolution and expands the receptive field. The feature map F2 enters the first DWABlock module (node ​​1-2), which efficiently extracts multi-scale information through a combination of deep convolution and point-by-point convolution to generate a feature map F3.

[0058] To further expand the receptive field and extract deep features, feature map F3 is then downsampled through a 3×3 convolution with a stride of 2 (nodes 1-3) to generate feature map F4. F4 enters another DWABlock module (nodes 1-4) to further extract multi-scale fusion features and output feature map F5. Next, F5 is downsampled again through a 3×3 convolution with a stride of 2 (nodes 1-5) to generate feature map F6. F6 continues to extract multi-scale features through another DWABlock module (nodes 1-6) to generate feature map F7. Next, F7 is further downsampled to feature map F8 through a 3×3 convolution with a stride of 2 (nodes 1-7), and enters the DWABlock module (nodes 1-8) to extract deep multi-scale features, generating feature map F9.

[0059] In the final stage of the feature extraction module, nodes 1-9 concatenate the feature map F7 of the middle layer with the feature map F9 of the deep layer to generate the fused feature map F10. Through cross-layer concatenation, the high-resolution features of the shallow layer and the low-resolution features of the deep layer are effectively combined, which not only retains the detailed information of the target, but also integrates the contextual semantic features, providing rich multi-scale feature expressions for subsequent modules.

[0060] 2. Upsampling and feature fusion model process based on nodes 1-10 to 1-22.

[0061] After feature extraction is completed, the network gradually restores the spatial resolution of the feature map through a series of upsampling and cross-layer feature fusion operations, while enhancing the detail information. First, the fused feature map F10 is input to the DWABlock module (node ​​1-10) to further optimize the feature expression and generate the feature map F11 for the ClsHead module (node ​​1-26) to perform the binary classification task of judging whether the image frame is abnormal based on the overall image frame in the first task branch.

[0062] In the second task branch, the feature map F7 is processed by convolution (nodes 1-11) and DWABlock module (nodes 1-12) in turn, and then upsampled by convolution transpose operation (ConvTranspose, nodes 1-13) to restore part of the spatial resolution and generate the upsampled feature map F12.

[0063] The feature map F12 is concatenated with the intermediate feature map F5 (from nodes 1-4) of the first stage in the Concat module (nodes 1-14), and the feature map F13 is generated by combining the shallow high-resolution features and the deep upsampled features. Subsequently, F13 is input into the DWABlock module (nodes 1-15) to perform multi-scale processing on the concatenated feature map to generate the feature map F14. F14 is further concatenated with the shallower feature map F3 (from nodes 1-2) in the Concat module (nodes 1-16) to form the feature map F15. Through this cross-layer fusion design, the network makes full use of the detailed information of the shallow features while combining the global semantic expression of the deep features.

[0064] In order to adjust the number of channels and resolution of the concatenated feature map, feature map F15 is processed by the DWABlock module (nodes 1-27) and a 3×3 convolution with a stride of 1 (nodes 1-17) to generate feature map F16. Subsequently, F16 is concatenated with F14 and further multi-scale features are extracted through the DWABlock module (nodes 1-19) to generate the optimized feature map F18. At this stage, the network fully combines deep low-resolution features with shallow high-resolution features through multiple upsampling and cross-layer feature concatenation operations, allowing the model to capture both the detailed information and contextual semantic relationships of the target.

[0065] 3. The second type of task branch based on nodes 1-19 to 1-25.

[0066] like Figure 2 As shown in the figure, the second task branch generates segmentation results of three different scales: large, medium and small through convolution, DWABlock module and fusion operation.

[0067] This multi-branch design can segment and optimize deep features and shallow features respectively, so that the segmentation results are balanced in detail preservation and global semantic expression. Especially in the fusion operation of nodes 1-18 and nodes 1-21, cross-layer splicing makes full use of the edge detail information in the shallow high-resolution features and the target semantic information in the deep low-resolution features, thereby significantly improving the segmentation accuracy.

[0068] exist Figure 2In the figure, the first and second task branches share part of the front network consisting of nodes 1-0 to 1-6. By sharing the feature-free extraction network, the two types of tasks can promote each other and further improve the overall performance of the model. It is particularly suitable for complex task scenarios that require simultaneous target recognition and segmentation, and provides an important guarantee for the model's multi-task learning ability.

[0069] exist Figure 2 In the OreSizeNet, multiple cross-layer feature fusion designs are adopted, so that shallow high-resolution features can be fully combined with deep low-resolution features in the decoding stage, which not only retains the detailed information of the target, but also integrates the global semantic expression. It is especially suitable for tasks with multi-scale targets, such as the segmentation of large pieces of ore and the fine classification of small pieces of ore. In the segmentation branch, the network fully refines and fuses the deep and shallow features through a multi-branch design, significantly improving the accuracy and robustness of the segmentation results.

[0070] In addition, the multiple applications of the DWABlock module have greatly improved the computational efficiency and feature extraction capabilities of the network. Through the efficient computation of separate convolutions and the channel fusion of point-by-point convolutions, the network can efficiently extract multi-scale features while maintaining a low computational cost. The upsampling stage gradually restores the resolution of the feature map and introduces shallow high-resolution features after each upsampling, allowing the network to maintain a balance between details and semantics when generating high-resolution outputs.

[0071] In this embodiment, DWABlock is a key module in OreSizeNet, which is used to efficiently extract multi-scale features while maintaining low computational complexity. Figure 1 As shown in the figure, the module combines a variety of depthwise separable convolutions (DWConv) of different sizes to fully model local and global information through multi-path feature extraction and fusion; the functions of each part are detailed as follows:

[0072] The input feature map first passes through a 1×1 standard convolution with a stride of 1 (node ​​2-1) to adjust the number of channels and reduce the amount of computation. The output of this step is the feature map T1, which serves as the basic feature input for subsequent operations.

[0073] The initial feature map T1 is copied to three parallel paths, and different convolution operations are performed on each path to achieve multi-path feature extraction:

[0074] A. Path 1 consisting of DWConv1×3 (node ​​2-2) and DWConv5×1 (node ​​2-3). In the first path, the feature map T1 first passes through a 1×3 depth-separable convolution (DWConv1×3) to extract local features in the horizontal direction. Then, the output feature passes through a 5×1 depth-separable convolution (DWConv5×1) to further capture vertical and larger context information.

[0075] B. Path 2 consisting of DWConv1×7 (nodes 2-3) and DWConv5×1 (nodes 2-4). In the second path, the feature map T1 passes through a 1×7 depthwise separable convolution (DWConv1×7) to extract features with a longer horizontal range. Then, the output features also pass through a 5×1 depthwise separable convolution to enhance the vertical detail information.

[0076] C. Path 3, which consists of DWConv5×1 (nodes 2-5) and DWConv1×7 (nodes 2-6), uses a 5×1 depthwise separable convolution on T1 to capture local features in the vertical direction. The output then passes through a 1×7 depthwise separable convolution to extract long-range information in the horizontal direction.

[0077] The output features of the three paths represent local and global multi-scale information respectively. Through the feature concatenation (Concat) operation (node ​​2-7), the features from different paths are combined into a multi-channel feature map T2. The concatenated multi-channel feature map T2 undergoes a 1×1 convolution operation (node ​​2-8) to compress and fuse the number of channels to generate the final output feature map T3 for use by subsequent modules.

[0078] In this embodiment, the DWABlock module can capture a variety of feature information from local details to long-distance global information by combining deep separable convolutions of different scales such as 1×3, 1×7, and 5×1, thereby improving the model's adaptability to multi-scale changes in the target. In addition, the deep separable convolution decomposes the standard convolution into deep convolution and point-by-point convolution, effectively reducing the number of parameters and computational overhead, so that DWABlock has good efficiency while maintaining high expressiveness. In this way, through path separation and cross-scale feature splicing, DWABlock achieves a good fusion between spatial features and contextual information, improving the integrity and semantic richness of feature expression.

[0079] like Figure 3As shown, the SegHead module of this embodiment is an efficient multi-task feature extraction and prediction structure, including three stages: initial feature adjustment, multi-layer feature extraction and fusion, and final task prediction, which is mainly used for ore block size measurement and anomaly identification. First, the input features are preliminarily adjusted through a 1×1 convolution (Conv1×1, node 3-1) to optimize the number of channels and computational complexity, laying the foundation for subsequent operations. Subsequently, after three layers of progressive feature extraction and fusion (DWABlock and Concat, nodes 3-2 to 3-7), multi-scale information is gradually captured and multi-level local and global features are integrated through feature splicing operations. After feature extraction at each stage, the features output by the DWABlock module are fused through the Concat operation to ensure full integration and transmission of information. Finally, the fused feature map is further adjusted through a 1×1 convolution (Conv1×1, node 3-8) to generate a unified high-dimensional feature representation to support multi-task prediction. The final output includes three types of prediction heads (nodes 3-9): classification (cls), rectangular box regression (rect), and segmentation mask (mask), which achieve accurate measurement of ore size and efficient identification of abnormal minerals. The modular design of SegHead effectively balances feature extraction efficiency and computational complexity in the size measurement task, and has significant practicality and robustness.

[0080] In this embodiment, Figure 4 As shown in the figure, the ClsHead module is an efficient feature extraction and prediction structure designed for classification tasks, which is mainly used for ore anomaly prediction and classification and identification of large ores. The overall process of the module includes three stages: initial feature adjustment, multi-path feature extraction and fusion, and final classification prediction. First, the input feature is adjusted through a 1×1 convolution (Conv1×1, node 4-1) to adjust the number of channels, optimize the computational complexity and lay the foundation for subsequent operations. Then, the adjusted features are sent to DWABlock (node ​​4-2), and the local abnormal features and global information of large ores are effectively captured through the multi-scale feature extraction mechanism. The features output by DWABlock are further fused with multi-path feature information through feature concatenation (Concat, node 4-3) to generate multi-dimensional feature expressions. Subsequently, the fused features are adjusted through a 1×1 convolution (Conv1×1, node 4-4) to compress redundant information and improve feature expression capabilities. Finally, after the classification head (cls, node 4-5), the module outputs the classification prediction results for large ores and abnormal ores. ClsHead effectively improves the prediction accuracy and robustness of ore classification and anomaly detection tasks through modular design and multi-scale feature fusion.

[0081] In this embodiment, in order to effectively optimize the ore size measurement and anomaly identification tasks, the following comprehensive loss function is used, which combines multiple loss components to simultaneously optimize mask segmentation, anomaly prediction and class imbalance problems. The comprehensive loss function consists of the following four parts:

[0082] 1. Binary Cross Entropy Loss (BCE): It is used to measure the difference between the model output and the true label. It is often used in binary classification tasks and has good numerical stability. Its expression is:

[0083] ;

[0084] in, Pixel The real label, is the predicted value (specifically a probability value between 0 and 1), is the total number of pixels.

[0085] 2. Dice loss: It is mainly used to solve the problem of category imbalance in segmentation tasks, especially when the target area is small. Dice loss improves the model's detection ability for small targets by optimizing the ratio of two overlapping areas. Its expression is:

[0086] ;

[0087] in, is a smoothing factor used to prevent the denominator from being zero.

[0088] 3. Tversky loss: Optimizes false negatives (FalseNegative, FN) and false positives (FalsePositive, FP) in abnormal prediction. Tversky loss introduces adjustment parameters and To control the penalty for false negatives and false positives to adapt to different types of anomaly detection tasks. Its expression is:

[0089] ;

[0090] in, and It is a hyperparameter used to balance the impact of false negatives and false positives and is usually adjusted according to the characteristics of the data.

[0091] 4. Weighted binary cross entropy loss: In order to deal with the problem of giving priority to large ore areas in ore size measurement, weighted binary cross entropy loss is used. By assigning different weights to different areas, the model can pay more attention to the prediction of large ore areas. Its expression is:

[0092] ;

[0093] Among them, the weight Dynamically adjust the pixel according to the area it belongs to. For example, for large ore areas, the weight Larger; for small pieces of ore or background areas, the weight Smaller.

[0094] To this end, the comprehensive loss function is the weighted sum of the above-mentioned losses, and its expression is:

[0095] ;

[0096] in, It is the weight hyperparameter of each part of the loss, which is used to balance the impact of different loss items on model training.

[0097] Through the design of the above-mentioned comprehensive loss function, we can effectively cope with multiple challenges in ore size measurement and anomaly identification, including category imbalance, small target segmentation, priority optimization of large ores, etc., thereby improving the segmentation accuracy and anomaly prediction ability of the model.

[0098]

OreSizeNet model training

[0099] The training process of the OreSizeNet model includes key steps such as data preparation, model initialization, loss function calculation, optimizer selection, and training strategy. The specific training process is as follows:

[0100] 1. Data preparation:

[0101] First, in order to ensure data diversity and sufficiency during training, a large-scale annotated dataset is used for training. The dataset includes ore images and their corresponding true labels (including blockiness, anomaly detection annotations, etc.). In the data preprocessing stage, we normalize, scale, and perform data enhancement on the images to increase the diversity of the dataset and improve the generalization ability of the model. Data enhancement techniques include random cropping, rotation, translation, flipping, etc.

[0102] 2. Model initialization:

[0103] During the training process, the weights of the pre-trained model are first loaded, or the model parameters are randomly initialized. For the OreSizeNet model, a backbone network with strong representation capabilities (such as ResNet, EfficientNet, etc.) is usually used for initialization. The feature extraction capabilities of these networks help to accelerate the training process and improve the accuracy of the model. Subsequently, all network layer parameters will be optimized through back propagation.

[0104] 3. Loss function:

[0105] During the training process, the comprehensive loss function designed above is used to optimize the model. The loss function includes binary cross entropy loss ( )、Dice loss( )、Tversky loss( ) and the weighted binary cross entropy loss ( ). These loss terms are weighted to take into account segmentation accuracy, class imbalance, and preferential attention to large blocks of ore during training.

[0106] Preferably, in the comprehensive loss function, The initial values ​​are: 0.4, 0.3, 0.3 and 0.1 respectively.

[0107] 4. Optimizer and learning rate adjustment:

[0108] During model training, optimization algorithms with adaptive learning rates, such as Adam or AdamW, are used to optimize model parameters. These optimizers can automatically adjust the learning rate according to the changes in the gradient, thereby accelerating the training process and improving convergence. In order to further improve the training effect, a learning rate decay strategy is adopted to gradually reduce the learning rate as the training iterations increase. This can prevent the model from oscillating in the later training process, thereby improving the accuracy of the final model. The goal of the optimization process is to minimize the loss function , calculate the gradient and update the network parameters through back propagation.

[0109] 5. Training strategy:

[0110] During the training process, the mini-batch stochastic gradient descent (SGD) method is used to update parameters. Each time during training, a mini-batch sample is randomly selected from the training data set, the loss function is calculated through forward propagation, and then the model parameters are updated through back propagation. Early stopping technology can also be used during training, that is, when the performance on the validation set no longer improves, training is stopped in advance to avoid overfitting.

[0111] In addition, in order to improve the generalization ability of the model, the cross-validation technique is used to effectively evaluate the performance of the model by dividing the dataset into multiple subsets for training and validation respectively.

[0112] 6. Model Evaluation:

[0113] During the training process, the performance of the model is monitored by regularly evaluating it on the validation set. Evaluation indicators include IoU (Intersection over Union), Dice coefficient, F1 score, precision, recall, etc. to comprehensively evaluate the performance of the model in blockiness measurement and anomaly identification tasks.

[0114] After training is completed, the final model will be evaluated for performance using a test set to ensure that it can stably and reliably perform ore size measurement and anomaly identification in practical applications.

[0115] 7. Hyperparameter Tuning:

[0116] During the training process, gradually adjust hyperparameters (such as learning rate, batch size, loss function weight, etc.) to optimize the performance of the model. Use hyperparameter optimization methods such as grid search or random search to find the optimal hyperparameter configuration.

[0117] 8. Training ends and model is saved:

[0118] After training is completed, save the final trained model, including the model weights, hyperparameter settings, and intermediate results during training. These saved model files can be directly used in subsequent reasoning and deployment stages.

[0119] [Calculation process of actual ore size]

[0120] Based on the mask segmentation results provided by the OreSizeNet model, assuming that the depth of the ore is known and fixed, the actual size of the ore can be accurately calculated through the camera's intrinsic and extrinsic calibration. The following is the detailed calculation process:

[0121] 1. Processing of segmentation results:

[0122] First, the OreSizeNet model is used to segment the ore image and obtain the mask of the ore area. The mask image accurately identifies the pixel area of ​​the ore and provides information such as the shape and outline of the ore. On this basis, the bounding box or minimum circumscribed polygon of the ore area can be extracted for subsequent size calculations.

[0123] 2. Camera internal and external parameters:

[0124] The camera's intrinsic and extrinsic parameters are used to convert image coordinates (pixel coordinates) into actual world coordinates. Specifically, the camera's intrinsic parameter matrix K describes the camera's focal length and the image's principal point coordinates, while the camera's extrinsic rotation matrix R and translation vector T describe the relationship between the camera coordinate system and the world coordinate system. For known camera intrinsic and extrinsic parameters, the coordinates of each pixel in the image are The three-dimensional coordinates can be converted into the actual world coordinate system through the following projection equation .

[0125] The camera intrinsic parameter matrix K is:

[0126] ;

[0127] in, and is the focal length, and The main point coordinates.

[0128] The camera extrinsic parameter is the rotation matrix and translation vectors , used to convert the three-dimensional point in the camera coordinate system Convert to a point in the world coordinate system :

[0129] .

[0130] 3. Conversion between pixel coordinates and actual coordinates:

[0131] At the assumed depth of the ore When fixed, the actual size calculation process of the ore can be simplified to only calculate the width in the image ( ) and height ( ). Assuming the depth is a known constant, then the camera's intrinsic matrix will help convert pixel coordinates to real coordinates.

[0132] For each pixel in the image , which corresponds to the actual size and It can be calculated by the following formula:

[0133] ;

[0134] ;

[0135] in, and are the width and height of the ore area in the image (in pixels), is the fixed depth of the ore (the actual depth in the world coordinate system), which is known or determined a priori.

[0136] 4. Calculation of actual size:

[0137] The actual size of the ore can be calculated by the pixel size of the bounding box or the minimum circumscribed polygon obtained from the mask segmentation result. The specific steps are as follows:

[0138] Bounding box extraction: Extract the bounding box of the ore area from the mask image to obtain the width of the ore in the image and height (pixel units).

[0139] Actual size calculation: using known depth And the camera internal parameters, convert the width and height of the pixel unit to the actual size and .

[0140] In addition, by measuring the key dimensions of the ore, such as length, width, and thickness, the volume of the ore can be calculated or other related geometric analysis can be performed.

[0141] 5. Final output:

[0142] Finally, through the above process, combined with the camera's internal and external calibration parameters and a known fixed depth, the model can output the actual size of the ore (such as width, height and depth), and provide important data information for ore block size measurement and anomaly identification. For practical applications, the size of the ore can be further refined for automated processing and analysis.

[0143] Usually, when the depth of the ore is known and fixed, based on the segmentation results of the OreSizeNet model and the calibration of the camera's internal and external parameters, the actual size of the ore can be directly calculated by converting pixel coordinates to actual coordinates. This process greatly simplifies the calculation process because only the width and height of the ore (in pixels) need to be considered and converted in combination with a fixed depth value. This method provides efficient and accurate support for tasks such as ore size measurement and anomaly identification.

[0144] Embodiment 2:

[0145] Corresponding to the above-mentioned embodiment, this embodiment discloses a system for measuring ore size and identifying anomalies, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the above-mentioned method is implemented when the processor executes the computer program.

[0146] In summary, the technical solutions disclosed in the above embodiments of the present invention respectively solve the technical problems and innovations mainly in the following aspects:

[0147] 1. High efficiency and accuracy of ore size measurement and anomaly identification.

[0148] The OreSizeNet network provided by the present invention realizes efficient processing of ore size measurement and anomaly identification tasks by combining multi-scale feature extraction, time series information fusion and deep separable convolution module (DWABlock). In a complex environment with a variety of ore sizes and shapes, the system can accurately identify ore size and effectively identify abnormal minerals. In particular, for the measurement of different ore sizes, the network can accurately handle the differences between large and small ore pieces, greatly improving the classification and segmentation accuracy.

[0149] 2. Multi-task learning design.

[0150] The present invention adopts a joint classification, segmentation and anomaly detection task model, and improves the multi-task learning ability of the model by sharing the feature extraction module. While measuring the ore block size, it can achieve accurate detection of abnormal ores, and the sharing mechanism of segmentation and classification tasks effectively reduces the number of parameters and improves the computing efficiency, so that the entire model shows excellent computing performance during training and reasoning.

[0151] 3. Efficient multi-scale feature extraction.

[0152] The DWABlock module, combined with multiple depth-separable convolutions, can efficiently extract multi-scale features in ore images while maintaining low computational complexity. This design can capture local detail information while combining global context information, improving the accuracy and robustness of ore block size measurement and anomaly identification. The DWABlock module also enhances the model's adaptability to ores of different scales through multi-path feature extraction.

[0153] 4. Optimized loss function design.

[0154] The comprehensive loss function proposed in this paper combines binary cross entropy loss, Dice loss, Tversky loss and weighted binary cross entropy loss, which can effectively solve the problems of category imbalance, small target segmentation, and large ore priority optimization in ore size measurement and anomaly recognition tasks. This loss function not only improves the segmentation accuracy, but also enhances the prediction ability of abnormal ores and adapts to complex ore scenarios.

[0155] 5. Modular design and efficient computing.

[0156] OreSizeNet adopts a modular design concept, and all modules are optimized to improve computing efficiency and feature extraction capabilities. Through layer-by-layer feature extraction, cross-layer feature fusion and upsampling operations, the network improves accuracy while ensuring efficient use of computing resources, allowing the system to ensure high-quality ore size measurement and anomaly identification results even with limited resources.

[0157] 6. High robustness and real-time performance.

[0158] The network design of the present invention enables the model to handle noise and changes in complex ore images, and maintain high robustness under different lighting, different backgrounds or different ore morphologies. Combined with the real-time processing requirements, the present invention can complete the ore size measurement and anomaly identification tasks in a relatively short time, meeting the real-time requirements on the production line.

[0159] 7. Wide applicability.

[0160] This technical solution is not only suitable for ore size measurement and abnormal ore identification, but can also be expanded to target segmentation and classification tasks in other complex scenarios. It has strong universality and flexibility.

[0161] In summary, the present invention successfully achieves the high efficiency, accuracy and real-time performance of ore size measurement and anomaly identification tasks through the innovative design of the OreSizeNet network structure, and has significant application value and technical advantages.

[0162] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for measuring ore size and identifying anomalies, characterized in that: include: Construct, train and verify the OreSizeNet network, which includes two types of task branches. The first type of task branch is a binary classification task based on whether the image frame is abnormal as a whole, and the second type of task branch is a segmentation task based on the pixel level of the image frame at three different scales: large, medium and small, for calculating the size of the ore. The first and second types of task branches share part of the front network. Based on the OreSizeNet network, the real-time collected images are synchronously measured for ore size and anomaly identification; The OreSizeNet network includes a standard convolution module for feature processing, a feature fusion module and a DWABlock module based on deep separable convolution; the DWABlock module is deployed in the front network shared by the first and second task branches; at least one DWABlock module is deployed in each intermediate network between the first and second task branches and before each classification head; and the DWABlock module is also deployed in each detection head; The DWABlock module first performs initial feature transformation with a 1×1 standard convolution unit with a stride of 1, and then extracts local and global multi-scale features with three parallel paths and then fuses them. Each path processes the horizontal and vertical information in series with two depth-wise separable convolutions. The first path is preceded and followed by two depth-wise separable convolutions of 1×3 and 5×1, respectively. The second path is preceded and followed by two depth-wise separable convolutions of 1×7 and 5×1, respectively. The third path is preceded and followed by two depth-wise separable convolutions of 5×1 and 1×7, respectively. The fused features are then output after the number of channels is adjusted by a 1×1 standard convolution. Among them, the classification heads of large, medium and small scales of the second task branch use the SegHead module. The data processing flow of the SegHead module includes three stages: initial feature adjustment, multi-layer feature extraction and fusion, and final task prediction, specifically: First, the input features are preliminarily adjusted through a 1×1 standard convolution, and then progressive feature extraction is performed using three combinations consisting of a DWABlock module and a feature fusion module. The first feature fusion module fuses the output of the first DWABlock module with the output of the preliminary adjustment standard convolution, the second feature fusion module fuses the output of the second DWABlock module with the output of the first DWABlock module, and the third feature fusion module fuses the output of the third DWABlock module with the output of the second DWABlock module. The output of the third feature fusion module is adjusted for the number of channels through a 1×1 standard convolution and then output for final classification prediction.

2. The method according to claim 1, characterized in that: Also includes: The actual size of the ore is calculated based on the segmentation results of the OreSizeNet network in the second task branch, the internal and external parameters of the camera calibration, and the conversion relationship from pixel coordinates to actual coordinates.

3. The method according to claim 1, characterized in that The classification head of the first type of task branch adopts the ClsHead module. In the ClsHead module, the input features are first adjusted to the number of channels through a 1×1 standard convolution, and the adjusted features are sent to the DWABlock module. The DWABlock output and the features of the standard convolution output with adjusted channel numbers are fused and then adjusted to the number of channels through a 1×1 standard convolution for final classification prediction.

4. The method according to any one of claims 1 to 3, characterized in that: The loss function of the OreSizeNet network manipulates the following four parts: Part 1: Binary Cross Entropy Loss for Measuring the Difference Between Model Output and True Label ; The calculation formula is: ; in, Pixel The real label, is the predicted value, is the total number of pixels; Part 2: Dice loss for solving category imbalance in segmentation tasks , the calculation formula is: ; in, is a smoothing factor used to prevent the denominator from being zero; The third part, Tversky loss, is calculated as follows: ; in, and is a hyperparameter; Part 4: Weighted Binary Cross Entropy Loss , the calculation formula is: ; Among them, the weight Dynamically adjust according to the area to which the pixel belongs; The final comprehensive loss function is the weighted sum of the above losses, and its expression is: ; in, is the weight hyperparameter of each part of the loss.

5. A system for measuring ore size and identifying anomalies, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Ore scale measurement method based on deep learning and application system

    CN110390691A

  • Deep learning ore size measurement method based on OfficientDet network and early warning system

    CN113158829A