Grinding rock target detection method and system based on HardRockNet network structure
By adopting the HardRockNet network structure of the stubborn stone target detection method during the grinding process, using multi-scale feature extraction and feature fusion technology, the stubborn stone detection accuracy, poor real-time performance and complex background problems are solved, and efficient and accurate stubborn stone detection and real-time optimization of the grinding process are achieved.
Patent Information
- Application Number
- CN202510181095.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-19
AI Technical Summary
The prior art has low accuracy and poor real-time detection of stubborn stones during grinding, and it is difficult to effectively deal with small goals, complex backgrounds and overlapping targets.
The grinding stone object detection method based on the HardRockNet network structure is adopted. By constructing a target detection model including a multi-scale feature extraction layer, a DFPN feature fusion layer and an output layer, the feature extraction is used for DPBlock module, the DFPN module performs feature fusion, and the target overlap problem is handled through the DetBlock module.
It significantly improves the detection accuracy of small targets, enhances the robustness of complex backgrounds and target overlap, realizes efficient real-time detection, and can adjust grinding process parameters in real time according to the detection results, improves ore processing efficiency and reduces costs.
Smart Images

Figure CN119672322B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of target detection, and in particular to a method and system for detecting grinding rock targets based on a HardRockNet (hard rock detection network) network structure. Background Art
[0002] Detection of stubborn rock in the grinding process is one of the key technologies for evaluating the difficult-to-grind properties of ores during ore processing. Ores with high stubborn rock content usually have poor grinding characteristics, resulting in low efficiency of the grinding process, increased energy consumption, and even affecting the quality of subsequent processing of the ore. Therefore, real-time detection of the presence and content of stubborn rock and adjustment of grinding process parameters based on the detection results have become the core tasks of optimizing the grinding process.
[0003] At present, the technology for detecting stubborn rocks in the grinding process is still in the research stage, and no relevant patent documents have been found. The detection of stubborn rocks faces several technical challenges, including small target size, complex background, and overlap between targets. Traditional stubborn rock detection methods usually rely on manual or traditional image processing technology-based algorithms, which often perform poorly in complex backgrounds and cannot effectively identify small or overlapping targets in real time.
[0004] With the development of computer vision and deep learning technology, object detection has become an important tool to solve this problem. In recent years, deep neural networks have made significant progress in the field of object detection, especially in small object detection and background interference suppression in complex scenes. Object detection frameworks based on convolutional neural networks (CNNs), such as SSD (Single Shot MultiBox Detector), YOLO (You Only Look Once) v5, and YOLOv8, have shown high detection accuracy and real-time performance in multiple applications. However, these classic object detection models still face the following limitations in the task of stubborn stone detection in the grinding process:
[0005] Insufficient small target detection capability:
[0006] SSD: Although SSD can generate multiple candidate boxes in a single forward propagation, its default basic network structure is relatively simple, and there is a problem of reduced accuracy when detecting smaller targets. Especially in the grinding process, the stubborn stone target is small and may overlap with other mineral particles. SSD performs poorly in detecting such small targets.
[0007] YOLOv5 and YOLOv8: The YOLO series performs well in real-time detection, but it also has limitations when dealing with small targets. In particular, the smaller receptive field design of YOLOv5 helps improve the detection accuracy of large targets, but it is more difficult to capture the boundaries of small targets. While YOLOv8 has made improvements in improving the accuracy of small target detection, it still has difficulty dealing with situations where it is difficult to separate the target from the background in complex backgrounds.
[0008] False detection problems caused by background complexity:
[0009] SSD: Although SSD uses a multi-scale feature fusion strategy to enhance the perception of targets of different sizes, it is prone to background misdetection when the target and background information are similar. Especially in the complex background of the grinding process, the texture and color of the ore particles are similar to the background, which increases the difficulty of distinguishing the target from the background.
[0010] YOLOv5 and YOLOv8: Although the YOLO series has strong real-time performance, its robustness in complex backgrounds is weak. YOLOv5 relies on fewer convolutional layers for feature extraction, so it is easy to ignore small target information in complex backgrounds, resulting in reduced target detection accuracy. YOLOv8 improves small target detection by optimizing the model architecture, but background interference may still affect the detection results in complex ore and dust backgrounds.
[0011] Target overlap and occlusion problems:
[0012] SSD: When faced with overlapping targets, SSD often causes detection confusion because multiple targets share similar areas. For target overlap that often occurs during the grinding process, SSD is prone to missing some targets, especially when the target is similar in color or shape to the background or other targets.
[0013] YOLOv5 and YOLOv8: The YOLO model also has the problem of poor detection of overlapping targets. Especially for the processing of small targets and overlapping targets, YOLOv5 optimizes this problem by assigning multiple prediction boxes, but it is still susceptible to target overlap. YOLOv8 improves this through more sophisticated anchor generation strategies and more complex training techniques, but it is still difficult to avoid misjudgment for dense target scenes. Summary of the invention
[0014] The invention provides a grinding rock target detection method and system based on a HardRockNet network structure, which are used to solve the technical problem of low accuracy of the grinding rock target detection method.
[0015] In order to solve the above technical problems, the technical solution proposed by the present invention is:
[0016] A method for detecting grinding rock targets based on HardRockNet network structure comprises the following steps:
[0017] Constructing a target detection model for grinding stones, the target detection model includes an input layer, a multi-scale feature extraction layer, a DFPN feature fusion layer and an output layer connected in series in sequence; the multi-scale feature extraction layer includes a plurality of feature extraction sublayers connected in series in sequence and having different feature extraction scales, the output ends of the plurality of feature extraction sublayers are all connected to the input end of the DFPN (Dual Feature Pyramid Network) feature fusion layer, and the plurality of feature extraction sublayers all use DPBlock modules for feature extraction;
[0018] Grinding pictures containing stubborn stones in different environments are collected, and the stubborn stones on the grinding pictures are marked to construct a training set, the training set is used to train the target detection model, and the trained target detection model is used to detect the stubborn stones on the grinding pictures.
[0019] Preferably, the feature extraction sublayer includes a downsampling module and a DPBlock (Dual-PathBlock) module connected in series, and the input layer is connected to the downsampling module at the head end of the multi-scale feature extraction layer; the output ends of the DPBlock modules are connected to the input ends of the DFPN feature fusion layer;
[0020] and / or
[0021] The multi-scale feature extraction layer includes a first feature extraction sublayer and a second feature extraction sublayer, the first feature extraction sublayer includes a first downsampling module and a first DPBlock module, the second feature extraction sublayer includes a second downsampling module and a second DPBlock module, the input end of the first downsampling module is connected to the output end of the input layer, the output end of the first downsampling module is connected to the input end of the first DPBlock module, the output end of the first DPBlock module is connected to the input end of the second downsampling module, the output end of the second downsampling module is connected to the input end of the second DPBlock module, and the output end of the first DPBlock module and the output end of the second DPBlock module are both connected to the input end of the DFPN feature fusion layer.
[0022] Preferably, the DPBlock module comprises: a global feature extraction unit, a local feature extraction unit and a fusion unit, and the output ends of the global feature extraction unit and the local feature extraction unit are both connected to the input end of the fusion unit.
[0023] Preferably, the global feature extraction unit comprises:
[0024] A first global feature extraction network, a first segmentation network, a first local feature extraction network, a first point-by-point convolution network, a first feature fusion network, a second point-by-point convolution network, and a first feature splicing network; the output end of the first global feature extraction network is connected to the input end of the first segmentation network, the output end of the first segmentation network is respectively connected to the input ends of the first local feature extraction network, the first point-by-point convolution network, and the first feature splicing network, the output end of the first local feature extraction network and the output end of the first point-by-point convolution network are both connected to the input end of the first feature fusion network, the output end of the first feature fusion network is connected to the input end of the first feature splicing network through the second point-by-point convolution network, and the output end of the first feature splicing network is connected to the input end of the fusion unit;
[0025] and / or
[0026] The local feature extraction unit comprises:
[0027] A second local feature extraction network, a second segmentation network, a second global feature extraction network, a third point-by-point convolutional network, a second feature fusion network, a fourth point-by-point convolutional network, and a second feature splicing network; the output end of the second local feature extraction network is connected to the input end of the second segmentation network, the output end of the second segmentation network is respectively connected to the input ends of the second global feature extraction network, the third point-by-point convolutional network, and the second feature splicing network, the output end of the second global feature extraction network and the output end of the third point-by-point convolutional network are both connected to the input end of the second feature fusion network, the output end of the second feature fusion network is connected to the input end of the second feature splicing network through the fourth point-by-point convolutional network, and the output end of the second feature splicing network is connected to the input end of the fusion unit;
[0028] and / or
[0029] The fusion unit includes a third feature fusion network and a fifth point-by-point convolutional network; the third feature fusion network is connected to the output ends of the global feature extraction unit and the local feature extraction unit respectively; the output end of the third feature fusion network is connected to the input end of the fifth point-by-point convolutional network, and the output end of the fifth point-by-point convolutional network is connected to the input end of the DFPN feature fusion layer.
[0030] Preferably, the DFPN feature fusion layer includes: multiple feature pyramid branches, the multiple feature pyramid branches correspond one-to-one to the multiple feature extraction sub-layers, each pyramid branch is used to extract the first feature of the output feature of the corresponding feature extraction sub-layer, and fuse and splice the extracted first feature with the first features extracted by other pyramid branches to output the corresponding fused feature.
[0031] Preferably, the DFPN feature fusion layer includes a first feature pyramid branch and a second feature pyramid branch;
[0032] The first feature pyramid branch includes a third DPBlock module, a fifth point-by-point convolutional network, a fourth feature fusion network, a fourth DPBlock module, a sixth point-by-point convolutional network and a third feature splicing network, the output end of the third DPBlock module is connected to the input end of the fourth feature fusion network through the fifth point-by-point convolutional network, the output end of the third DPBlock module is also directly connected to the input end of the fourth feature fusion network, the output end of the fourth feature fusion network is connected to the input end of the fourth DPBlock module, the output end of the fourth DPBlock module is connected to the input end of the third feature splicing network through the sixth point-by-point convolutional network, and the output end of the fourth DPBlock module is also directly connected to the input end of the third feature splicing network;
[0033] The second feature pyramid branch includes a fifth DPBlock module, a seventh point-by-point convolutional network, a fourth feature splicing network, a sixth DPBlock module, an eighth point-by-point convolutional network and a fifth feature fusion network, the output end of the fifth DPBlock module is connected to the input end of the fourth feature splicing network through the seventh point-by-point convolutional network, the output end of the fifth DPBlock module is also directly connected to the input end of the fourth feature splicing network, the output end of the fourth feature splicing network is connected to the input end of the sixth DPBlock module, the output end of the sixth DPBlock module is connected to the input end of the fifth feature fusion network through the eighth point-by-point convolutional network, and the output end of the sixth DPBlock module is also directly connected to the input end of the fifth feature fusion network;
[0034] The third DPBlock module is jump-connected to the fourth feature splicing network, the fifth DPBlock module is jump-connected to the fourth feature fusion network, the fourth DPBlock module is jump-connected to the fifth feature fusion network, and the sixth DPBlock module is jump-connected to the third feature splicing network.
[0035] Preferably, the output layer includes: a plurality of DetBlock modules, the plurality of DetBlock modules correspond one-to-one to the plurality of feature pyramid branches, and each DetBlock module performs stone target detection based on fused features outputted from the corresponding feature pyramid branch.
[0036] Preferably, the DetBlock module (Detection Block) includes: a first channel adjustment convolution layer, a seventh DPBlock module, a fifth feature splicing network, a second channel adjustment convolution layer, an eighth DPBlock module, a sixth feature splicing network and a third channel adjustment convolution layer;
[0037] The output end of the first channel adjustment convolution layer is connected to the input end of the fifth feature splicing network through the seventh DPBlock module, the output end of the first channel adjustment convolution layer is also directly connected to the input end of the fifth feature splicing network, the output end of the fifth feature splicing network is connected to the input end of the second channel adjustment convolution layer, the output end of the second channel adjustment convolution layer is connected to the input end of the sixth feature splicing network through the eighth DPBlock module, the output end of the second channel adjustment convolution layer is also directly connected to the input end of the sixth feature splicing network, and the output end of the sixth feature splicing network is connected to the input end of the third channel adjustment convolution layer.
[0038] Preferably, the loss function of the target detection model satisfies:
[0039] ;
[0040] ;
[0041] ;
[0042] ;
[0043] in, is the classification loss, is the bounding box regression loss, Optimize loss for background separation, , , is the importance weight to balance each loss item, is the predicted probability of the target category; is the category balancing factor; It is a tuning parameter that controls the weights of easy-to-classify samples and hard-to-classify samples; It is the intersection of the predicted box and the true box; Represents the Euclidean distance between the center points of the predicted box and the true box; Represents the coordinates of the center point of the prediction box, Indicates the coordinates of the center point of the real frame;
[0044] is the diagonal length of the bounding box; is a shape consistency measure, is the equilibrium parameter; The true label representing the background or target; is the probability predicted by the model; is the total number of samples.
[0045] A computer system comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the steps of the above method are implemented when the processor executes the computer program.
[0046] The present invention has the following beneficial effects:
[0047] 1. The present invention adopts the HardRockNet network structure and specially designs the DPBlock module, which effectively improves the detection ability of small targets through multi-scale feature extraction and deep convolution operations. The module can capture the local fine-grained features of small targets, thereby significantly improving the detection ability of small changes in stubborn stone targets in the ore background; in the complex ore background, even if the stubborn stone target is small, HardRockNet can still achieve high detection accuracy and reduce missed detection. In addition, the HardRockNet network structure in the present invention can effectively fuse multi-scale features and interactively enhance contextual information through the DFPN module. This design effectively suppresses the interference of background noise on target detection, especially in complex grinding scenarios, and can effectively reduce the problem of false detection caused by background complexity. Especially in the overlapping area of ore and other minerals, DFPN enhances the model's perception of the target and improves its robustness to complex backgrounds. Furthermore, the design of HardRockNet has high computational efficiency and can quickly process image data and make predictions in a grinding environment with high real-time requirements. The network greatly improves the speed of target detection through efficient feature extraction and multi-scale fusion, providing strong support for real-time adjustment of the grinding process. Using this method, the system can detect the distribution of stubborn stones in real time during the grinding process, and adjust the process parameters according to the detection results, thereby improving the processing efficiency of the ore and reducing the grinding cost.
[0048] 2. In the preferred embodiment, the present invention uses the DetBlock module to adopt the method of step-by-step feature fusion and spatial information enhancement to particularly effectively solve the problem of target overlap in the grinding process. DetBlock can refine the feature expression of the target, so that the network can distinguish overlapping targets and avoid the common misjudgment phenomenon in traditional methods. This feature is particularly suitable for the overlapping detection of a large number of crushed stone and stubborn stone targets in the grinding process, ensuring the accurate identification and positioning of overlapping targets.
[0049] 3. In the preferred embodiment, the present invention adopts a multi-task joint loss function (MultiTask Loss) to optimize the classification loss, bounding box regression loss and background suppression loss, so that the network can not only improve the classification accuracy of the target, but also accurately regress the bounding box position of the target, while reducing the confusion between complex background and small targets, further improving the detection accuracy. Especially for small targets and overlapping targets in complex scenes, this loss function can help the network better locate and distinguish different targets, and improve the comprehensiveness and accuracy of detection.
[0050] In addition to the above-described purposes, features and advantages, the present invention has other purposes, features and advantages. The present invention will be further described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] The drawings constituting a part of this application are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0052] Figure 1 A structural block diagram of a target detection model provided by an embodiment of the present invention;
[0053] Figure 2 A structural diagram of a DPBlock module provided in an embodiment of the present invention;
[0054] Figure 3 is a structural diagram of a DFPN module provided in an embodiment of the present invention;
[0055] Figure 4 is a structural diagram of a DetBlock module provided in an embodiment of the present invention; DETAILED DESCRIPTION
[0056] The embodiments of the present invention are described in detail below with reference to the accompanying drawings, but the present invention can be implemented in many different ways as defined and covered by the claims.
[0057] The purpose of the present invention is to solve the problems of low precision, poor real-time performance and insufficient detection capability of complex background and small targets in the existing technology during the grinding process. Specifically, the existing methods for detecting stubborn rocks usually face the following technical challenges:
[0058] 1. Small targets of stubborn stones: In the grinding process, stubborn stones usually appear as small targets, which makes them easily submerged by background noise in the ore. Traditional target detection methods are difficult to effectively capture the characteristics of these small targets, resulting in low detection accuracy.
[0059] 2. Complex background interference: During the grinding process, the complex background of the ore and the mixture of various minerals make it easy for the stubborn stone target to have similar features with the background area. Traditional target detection models perform poorly in these complex scenarios and are prone to false detection and missed detection.
[0060] 3. Target overlap problem: During the grinding process, the stubborn stone target sometimes overlaps with other targets (such as ore fragments), resulting in misjudgment and inability to accurately separate the targets during the detection process. Traditional methods often cannot effectively handle target overlap, affecting detection accuracy.
[0061] 4. Real-time detection and process adjustment requirements: The detection of stubborn stones in the grinding process needs to be efficient and real-time, so that the grinding process parameters can be adjusted in time according to the detection results, thereby optimizing the processing efficiency of the ore. However, existing technologies cannot meet the real-time requirements while ensuring high accuracy.
[0062] Therefore, the present invention proposes a grinding rock target detection method based on the HardRockNet network structure, aiming to solve the above problems, improve the small target detection accuracy through an improved deep learning model, enhance the robustness to complex background and target overlapping problems, and finally realize the accurate real-time detection of rock in the grinding process, thereby providing effective support for the optimization and real-time adjustment of the grinding process.
[0063] Specifically, the grinding rock target detection method based on the HardRockNet network structure in the present invention comprises the following steps:
[0064] 1. Building a target detection model for grinding rock (also known as HardRockNet network)
[0065] like Figure 1 As shown in the figure, the HardRockNet network mainly includes: an input layer, a multi-scale feature extraction layer, a DFPN feature fusion layer and an output layer connected in series in sequence; the multi-scale feature extraction layer includes multiple feature extraction sublayers connected in series in sequence and with different feature extraction scales, the output ends of the multiple feature extraction sublayers are all connected to the input end of the DFPN feature fusion layer, and the multiple feature extraction sublayers all use the DPBlock module for feature extraction.
[0066] In a preferred embodiment, the feature extraction sublayer includes a downsampling module and a DPBlock module connected in series, the input layer is connected to the downsampling module at the head end of the multi-scale feature extraction layer; the output ends of the DPBlock modules are connected to the input ends of the DFPN feature fusion layer;
[0067] In a preferred embodiment, the multi-scale feature extraction layer includes a first feature extraction sublayer and a second feature extraction sublayer, the first feature extraction sublayer includes a first downsampling module and a first DPBlock module, the second feature extraction sublayer includes a second downsampling module and a second DPBlock module, the input end of the first downsampling module is connected to the output end of the input layer, the output end of the first downsampling module is connected to the input end of the first DPBlock module, the output end of the first DPBlock module is connected to the input end of the second downsampling module, the output end of the second downsampling module is connected to the input end of the second DPBlock module, and the output end of the first DPBlock module and the output end of the second DPBlock module are both connected to the input end of the DFPN feature fusion layer.
[0068] In a preferred embodiment, the input layer is a 3×3 convolution block with a step size of 1, the first downsampling module and the second downsampling module are 3×3 convolution blocks with a step size of 2; the output layer is a plurality of DPBlock modules;
[0069] In this embodiment, the workflow of the HardRockNet network is as follows:
[0070] 1. The input features are first passed through a series of 3×3 convolution operations with strides of 1 and 2 for basic feature extraction.
[0071] 2. The extracted features are sent to multiple DPBlock modules to enhance the local feature expression capabilities layer by layer.
[0072] 3. The multi-scale features output by the DPBlock module are converged into the DFPN module, and the perception ability of small targets is further improved through bidirectional interactive feature fusion.
[0073] 4. Finally, the features are processed by two parallel DetBlock modules to generate detection results, including the confidence score (Score) and bounding box (Rect) of the target.
[0074] Among them, the DPBlock module is a module designed specifically for improving small target detection capabilities. Its core idea is to effectively capture the local information and fine-grained features of the target through multi-scale feature fusion and deep convolution operations. The module takes the input tensor as the starting point and performs multi-scale feature extraction and fusion through two parallel branches.
[0075] The DPBlock module is the core module of HardRockNet, which is designed to improve the detection capability of small targets. Its core idea is to effectively capture the local information and fine-grained features of the target through multi-scale feature fusion and deep convolution operations. Starting from the input tensor, DPBlock uses two parallel branches to extract features respectively:
[0076] Main branch: Use standard convolution operations to enhance the ability to express global features.
[0077] Auxiliary branch: The receptive field is expanded by dilated convolution operations to capture more contextual information.
[0078] The features of the two branches are fused at the end of the module to achieve effective expression of small target features while enhancing the robustness to complex background interference.
[0079] Specifically, Figure 2 As shown, the DPBlock module includes a global feature extraction unit, a local feature extraction unit and a fusion unit, and the output ends of the global feature extraction unit and the local feature extraction unit are both connected to the input end of the fusion unit.
[0080] Wherein, the global feature extraction unit comprises:
[0081] A first global feature extraction network, a first segmentation network, a first local feature extraction network, a first point-by-point convolution network, a first feature fusion network, a second point-by-point convolution network, and a first feature splicing network; the output end of the first global feature extraction network is connected to the input end of the first segmentation network, the output end of the first segmentation network is respectively connected to the input ends of the first local feature extraction network, the first point-by-point convolution network, and the first feature splicing network, the output end of the first local feature extraction network and the output end of the first point-by-point convolution network are both connected to the input end of the first feature fusion network, the output end of the first feature fusion network is connected to the input end of the first feature splicing network through the second point-by-point convolution network, and the output end of the first feature splicing network is connected to the input end of the fusion unit;
[0082] The local feature extraction unit comprises:
[0083] A second local feature extraction network, a second segmentation network, a second global feature extraction network, a third point-by-point convolutional network, a second feature fusion network, a fourth point-by-point convolutional network, and a second feature splicing network; the output end of the second local feature extraction network is connected to the input end of the second segmentation network, the output end of the second segmentation network is respectively connected to the input ends of the second global feature extraction network, the third point-by-point convolutional network, and the second feature splicing network, the output end of the second global feature extraction network and the output end of the third point-by-point convolutional network are both connected to the input end of the second feature fusion network, the output end of the second feature fusion network is connected to the input end of the second feature splicing network through the fourth point-by-point convolutional network, and the output end of the second feature splicing network is connected to the input end of the fusion unit;
[0084] The fusion unit includes a third feature fusion network and a fifth point-by-point convolutional network; the third feature fusion network is connected to the output ends of the global feature extraction unit and the local feature extraction unit respectively; the output end of the third feature fusion network is connected to the input end of the fifth point-by-point convolutional network, and the output end of the fifth point-by-point convolutional network is connected to the input end of the DFPN feature fusion layer.
[0085] Among them, in the preferred scheme, the first global feature extraction network and the second global feature extraction network are both 5×5 deep convolution (DWConv5X5), the first local feature extraction network and the second local feature extraction network are both 3×3 deep convolution (DWConv3X3); the first segmentation network and the second segmentation network are both split segmentation networks, and the first point-by-point convolution network to the fifth point-by-point convolution network are all Conv1X1 convolution; the first feature fusion network, the second feature fusion network, and the third feature fusion network all use element-by-element addition (ADD), and the first feature splicing network and the second feature splicing network are both feature splicing (Concat) networks.
[0086] like Figure 2 As shown, the workflow of the DPBlock module is as follows:
[0087] The input of the DPBlock module is a multi-channel feature tensor. First, local features are extracted through two different deep convolution operations. In branch 1, the input tensor is fed into a 3×3 deep convolution (DWConv3X3), which effectively extracts relatively small receptive field information to capture local detail features of the target. Branch 2 uses a 5×5 deep convolution (DWConv5X5) to focus on a larger range of contextual information with a larger receptive field. This design takes into account the needs of both local and global feature extraction, and provides strong support for distinguishing small targets from background information.
[0088] The output feature maps of the two branches are equally divided into three parts in the channel dimension (each part is one-third of the input channel), thereby achieving fine-grained feature decomposition. In branch 1, one-third of the feature map is fed into a 5×5 deep convolution (DWConv5X5) to further enhance the local feature expression capability, while the other one-third of the feature map is subjected to 1×1 point-by-point convolution (Conv1X1) for inter-channel feature interaction. This operation can improve the semantic expression capability of the features while maintaining the integrity of local features. The outputs of the two operations are fused through element-by-element addition (ADD), and finally a 1×1 point-by-point convolution (Conv1X1) is used to further compress the feature dimension to ensure a balance between computational efficiency and expression capability.
[0089] In branch 2, similar feature decomposition operations are applied: one-third of the feature map is processed by a 3×3 depthwise convolution (DWConv3X3) to extract finer-grained local features, while the other one-third of the feature map is processed by a 1×1 point-by-point convolution (Conv1X1). After fusing the two parts of the features through element-by-element addition (ADD), the features are compressed and integrated again through a 1×1 point-by-point convolution (Conv1X1). This gradual refinement process helps to extract salient features of small objects and reduce noise interference.
[0090] The output features of branch 1 and branch 2 are concatenated in the channel dimension to generate two multi-scale feature maps. The two feature maps are fused by element-wise addition (ADD), which effectively integrates the context information and local features of different branches. Finally, the fused feature map is dimensionally compressed by a 1×1 point-by-point convolution (Conv1X1) to generate the final output tensor.
[0091] The DPBlock module effectively enhances the detection capability of small targets through multi-scale deep convolution operations and fine-grained feature decomposition. First, deep convolution (DWConv) focuses on local features on a smaller receptive field to ensure that the boundary information and texture features of small targets can be accurately captured. Secondly, the combination of feature decomposition and point-by-point convolution (Conv1X1) achieves semantic enhancement between channels and improves the distinguishability of feature expression. Finally, the fusion of branch features and the integration of contextual information further eliminate the interference of background noise, ensuring that small targets still have high detection accuracy in complex scenes.
[0092] In general, DPBlock provides strong support for small object detection by combining multi-scale feature extraction, inter-channel feature interaction and context information integration. In particular, DPBlock demonstrates excellent robustness and detection capabilities in scenes with high background complexity and dense distribution of small objects.
[0093] Specifically, in order to solve the problem of information loss in multi-scale target (especially small target) detection, HardRockNet introduces the DFPN feature fusion layer as the core unit of multi-scale feature fusion. The DFPN feature fusion layer improves the network's detection ability for multi-scale targets through bidirectional interactive connections and hierarchical feature aggregation. The structure of DFPN is as follows Figure 3 As shown, it consists of two main feature pyramid branches (P1 and P2) and their interactive connections.
[0094] Two feature pyramid branches (P1 and P2): Each branch processes features of different scales and generates multi-scale feature representations through convolution operations of different step lengths.
[0095] Interactive connection: Through the interactive fusion of upper and lower layer features, the one-way connection limitation of the traditional feature pyramid is broken, and the expression ability of multi-scale features is further enhanced.
[0096] The DFPN feature fusion layer can effectively enhance the robustness of the network in complex scenarios, especially for the detection of small targets.
[0097] In a preferred embodiment, the DFPN feature fusion layer includes: multiple feature pyramid branches, each of which corresponds one-to-one to the multiple feature extraction sub-layers, and each pyramid branch is used to extract the first feature of the output feature of the corresponding feature extraction sub-layer, and fuse and splice the first feature extracted by it with the first features extracted by other pyramid branches to output the corresponding fused feature.
[0098] In a preferred embodiment, the DFPN feature fusion layer includes a first feature pyramid branch and a second feature pyramid branch;
[0099] The first feature pyramid branch includes a third DPBlock module, a fifth point-by-point convolutional network, a fourth feature fusion network, a fourth DPBlock module, a sixth point-by-point convolutional network and a third feature splicing network, the output end of the third DPBlock module is connected to the input end of the fourth feature fusion network through the fifth point-by-point convolutional network, the output end of the third DPBlock module is also directly connected to the input end of the fourth feature fusion network, the output end of the fourth feature fusion network is connected to the input end of the fourth DPBlock module, the output end of the fourth DPBlock module is connected to the input end of the third feature splicing network through the sixth point-by-point convolutional network, and the output end of the fourth DPBlock module is also directly connected to the input end of the third feature splicing network;
[0100] The second feature pyramid branch includes a fifth DPBlock module, a seventh point-by-point convolutional network, a fourth feature splicing network, a sixth DPBlock module, an eighth point-by-point convolutional network and a fifth feature fusion network, the output end of the fifth DPBlock module is connected to the input end of the fourth feature splicing network through the seventh point-by-point convolutional network, the output end of the fifth DPBlock module is also directly connected to the input end of the fourth feature splicing network, the output end of the fourth feature splicing network is connected to the input end of the sixth DPBlock module, the output end of the sixth DPBlock module is connected to the input end of the fifth feature fusion network through the eighth point-by-point convolutional network, and the output end of the sixth DPBlock module is also directly connected to the input end of the fifth feature fusion network;
[0101] The third DPBlock module is jump-connected to the fourth feature splicing network, the fifth DPBlock module is jump-connected to the fourth feature fusion network, the fourth DPBlock module is jump-connected to the fifth feature fusion network, and the sixth DPBlock module is jump-connected to the third feature splicing network.
[0102] In a preferred embodiment, the fifth point-by-point convolutional network to the eighth point-by-point convolutional network are all 1×1 point-by-point convolutions (Conv1X1);
[0103] The fourth feature fusion network and the fifth feature fusion network are both ADD networks, and the third feature concatenation network and the fourth feature concatenation network are both concatenation (Concat) networks.
[0104] like Figure 3 As shown, the DFPN feature fusion layer workflow is as follows:
[0105] 1. Input features and preliminary processing
[0106] The DFPN feature fusion layer receives input tensors from feature layers of different scales, which are recorded as and .
[0107] for :The input tensor is first processed by a DPBlock module. DPBlock enhances the ability to express local details and small object features through a multi-scale feature extraction mechanism. Subsequently, it passes through a 1×1 point-by-point convolution (Conv1X1) to further compress the channel dimension and extract deep semantic information.
[0108] for : The same operation is applied to the input tensor, extracting features through a DPBlock module and a 1×1 point-by-point convolution module.
[0109] 2. Feature interaction and fusion
[0110] After completing the initial feature extraction, and The features are exchanged through interactive connections to achieve complementary fusion of upper and lower layer features.
[0111] From the high-level features ( ) is passed to the lower-level features ( ): Through a cross-layer connection, the information of high-level features is passed to the low-level feature branch, strengthening the expression of global context information of low-level features.
[0112] From the low-level features ( ) is passed to high-level features ( ): Similarly, low-level features are passed to high-level branches through interactive connections, providing more fine-grained information for high-level features.
[0113] After feature interaction, In the branch, feature maps are fused by element-wise addition (ADD). In the branch, feature maps are fused by concatenation. This differentiated feature fusion strategy can capture more detailed information and semantic associations in feature maps of different scales.
[0114] 3. Multi-level feature extraction
[0115] The fused feature maps are passed through a DPBlock module and a 1×1 point-by-point convolution operation to further refine the feature expression capability. For each branch, the processed feature map contains richer contextual information and semantic features.
[0116] 4. Output and multi-scale object detection
[0117] High-level output ( ): The final features of the branch are concatenated to generate a high-level output feature map. This feature map has strong semantic information and is suitable for detecting large targets.
[0118] Low-level output ( ): The final features of the branch are generated through element-wise addition (ADD) operations to generate low-level output feature maps. This feature map has strong local detail information and is particularly suitable for detecting small targets.
[0119] In addition, the DFPN feature fusion layer also performs downsampling and upsampling operations (such as Figure 3As shown in Figure 3, the information transmission and complementation between features of different scales are further strengthened, ensuring that multi-scale targets can be accurately detected.
[0120] The design of the DFPN feature fusion layer significantly improves the detection performance of multi-scale objects through bidirectional feature interaction and differentiated fusion strategies.
[0121] The multi-scale extraction capability of the DPBlock module enhances the network's detection effect on small targets.
[0122] The interactive connection between high-level and low-level features makes up for the shortcomings of a single feature layer, allowing the global context information and local detail information to be effectively integrated.
[0123] By combining element-by-element addition and concatenation strategies, the expressiveness of features at different scales is further optimized.
[0124] In summary, the DFPN feature fusion layer successfully solves the problems of insufficient context and inaccurate small target detection in traditional feature fusion through the design of a bidirectional feature pyramid, and provides an efficient multi-scale feature processing solution for target detection tasks.
[0125] The DetBlock module is specifically designed for terminal processing of rock target detection. Its multi-stage architecture effectively improves the network's ability to locate and classify targets through feature extraction, feature fusion, and spatial information enhancement. The design of DetBlock can solve the problem of target overlap in rock detection while aggregating the feature details of small targets.
[0126] In this embodiment, if Figure 4 As shown, the DetBlock module includes: a first channel adjustment convolution layer, a seventh DPBlock module, a fifth feature splicing network, a second channel adjustment convolution layer, an eighth DPBlock module, a sixth feature splicing network, and a third channel adjustment convolution layer;
[0127] The output end of the first channel adjustment convolution layer is connected to the input end of the fifth feature splicing network through the seventh DPBlock module, the output end of the first channel adjustment convolution layer is also directly connected to the input end of the fifth feature splicing network, the output end of the fifth feature splicing network is connected to the input end of the second channel adjustment convolution layer, the output end of the second channel adjustment convolution layer is connected to the input end of the sixth feature splicing network through the eighth DPBlock module, the output end of the second channel adjustment convolution layer is also directly connected to the input end of the sixth feature splicing network, and the output end of the sixth feature splicing network is connected to the input end of the third channel adjustment convolution layer.
[0128] Among them, the first channel adjustment convolution layer to the third channel adjustment convolution layer are all 1×1 convolutions, and the fifth feature splicing network and the sixth feature splicing network are both feature splicing (Concat) networks;
[0129] Among them, the workflow of the DetBlock module is as follows:
[0130] First, the input features are processed by 1×1 convolution to reduce the channel dimension to reduce the computational burden while retaining the spatial resolution, laying the foundation for subsequent small target feature extraction.
[0131] The features after the initial 1×1 convolution enter the dual-path block (DPBlock). The DPBlock module captures multi-scale context information and local spatial detail features through parallel path design, thereby enhancing sensitivity to small targets. In particular, the DPBlock module can effectively suppress background noise and highlight the feature differences of overlapping targets.
[0132] The output features of the DPBlock module are concatenated with the initial input features through skip connections, fusing low-level detail information with high-level semantic information. This multi-level feature fusion is particularly important for small target detection, as it can improve the problem of feature ambiguity caused by the small size of small targets.
[0133] The concatenated features are further integrated through another 1×1 convolutional layer to improve the compactness of feature representation and enhance the ability to recognize subtle differences between objects.
[0134] The features are passed to the second DPBlock module again to further enhance the multi-scale feature expression, especially the feature distinction ability of small objects and overlapping objects. Subsequently, the output features are concatenated with the input features of this stage to fully integrate the deep semantic information with the shallow detail information.
[0135] After the final 1×1 convolutional layer, the DetBlock module generates an output feature map, providing refined feature representation to ensure the detection accuracy of small targets while alleviating the problem of false detection caused by target overlap.
[0136] In the task of detecting stubborn stone objects, the DetBlock module specifically addresses the two major challenges of small object detection and object overlap, significantly improving detection performance through multi-scale feature aggregation and multi-level information fusion. The final output includes the target's Score (confidence score) and Rect (detection box coordinates), providing strong support for accurately locating and classifying stubborn stone objects.
[0137] 2. Collect grinding pictures containing stubborn stones in different environments, and mark the stubborn stones on the grinding pictures to build a training set
[0138] 1. Data collection and annotation
[0139] In the task of detecting stubborn rocks, data collection and annotation are the basis of model training and an important part of improving detection accuracy. In order to ensure the quality of data, data collection needs to cover as many diverse scenes as possible. First, during the data collection stage, high-resolution cameras or industrial detection equipment should be used to take pictures of stubborn rocks. These pictures should include different background scenes (such as smooth backgrounds, complex texture backgrounds, backgrounds of different materials), various lighting conditions (such as bright, dim, shadows, reflections, etc.), and occlusion scenes (such as cases where the target is partially covered or overlapped). In this way, it can be ensured that the dataset contains diverse samples that are consistent with the actual detection scene. At the same time, in order to capture the detailed features of the target, samples containing small targets should be collected. In addition, to ensure sufficient data volume, at least thousands of pictures should be collected for each type of target.
[0140] After completing the data collection, the next step is data labeling. Use labeling tools (such as LabelImg) to manually label the images and generate the bounding box (Bounding Box) and category label of the target. When labeling the bounding box, make sure that the box tightly surrounds the target and avoid including redundant background information. In particular, the labeling of small targets must be extremely precise to ensure that no details are missed. In addition, for overlapping targets, each target should be labeled separately to avoid labeling conflicts. In terms of category labeling, the target classification label should be determined according to the task requirements. If the task only involves target detection, it can be uniformly labeled as one category; if the task needs to distinguish between target types, a unique label needs to be assigned to each category. In order to ensure the high quality and consistency of the labeled data, all labelers should follow the unified labeling specifications and review and verify the labeling results multiple times to minimize labeling errors.
[0141] 2. Data Preprocessing
[0142] After completing data collection and annotation, the data needs to be preprocessed to further improve the generalization ability of the model. The first step of data preprocessing is to apply image enhancement technology to expand the diversity of data. Geometric transformation is one of the commonly used enhancement methods, including random cropping, rotation, and flipping. For example, by randomly cropping different areas of the target and scaling them back to the original size, the different positions of the target in the image can be simulated; by rotating and flipping, the directional diversity of the target can be increased. In addition, color perturbations such as brightness adjustment, contrast adjustment, and saturation adjustment can be used to simulate target changes under complex lighting conditions. In order to further improve the robustness of the model, Gaussian noise or salt and pepper noise can also be added to simulate interference in the environment. Finally, in order to enhance the model's ability to detect partially occluded targets, part of the image can be randomly occluded to simulate occlusions that may occur in actual scenes.
[0143] In the second step of preprocessing, a multi-scale training strategy is used to improve the model's adaptability to objects of different sizes. Specifically, the images are randomly scaled to different resolutions (such as 320×320, 512×512, 640×640) and then input into the model, so that the model can learn the features of large and small objects at the same time. This multi-scale training method is particularly suitable for small object detection tasks and can significantly improve the recall rate and detection accuracy of the model.
[0144] The last step of data preprocessing is data partitioning. In order to ensure the scientificity and fairness of model training, the labeled data set should be divided into training set, validation set and test set in a ratio of 7:2:1. The training set is used for model parameter learning, the validation set is used for performance evaluation during model training (such as detecting overfitting), and the test set is used for performance verification of the final model. When partitioning the data, samples need to be randomly selected while ensuring that each category is evenly distributed in the training set, validation set and test set. It should be noted that data augmentation is only applied to the training set, and the validation set and test set should keep the original data to avoid the authenticity of the evaluation results being affected by the augmentation.
[0145] Through the above detailed data collection, labeling and preprocessing steps, a high-quality data foundation can be built for the training of the HardRockNet model. This can not only improve the model's detection ability for small and overlapping targets, but also enhance the model's robustness and generalization ability in complex backgrounds, laying the foundation for the success of the hard rock target detection task.
[0146] 3. Use the training set to train the target detection model.
[0147] 1. Pre-trained weight loading
[0148] In deep learning model training, loading pre-trained weights is an effective way to improve the initial performance of the model, especially when the amount of training data is limited. For the HardRockNet model, if there are pre-trained weights available (such as models based on ImageNet), they can be applied to the backbone network of the model. These weights have learned basic image features (such as edges, textures, etc.), providing a good starting point for the model, thereby accelerating convergence and improving detection performance. Pre-trained weights can be loaded through built-in functions of common deep learning frameworks (such as PyTorch or TensorFlow). For example, for ResNet or other mainstream backbone networks, the official pre-trained model can be called directly.
[0149] For HardRockNet's dedicated modules (such as DPBlock and DFPN), since they are custom-designed modules, they usually do not have pre-trained weights. At this time, the weights of these modules need to be properly initialized. Common methods include Xavier initialization and He initialization. The former is suitable for general network layers, and the latter is particularly suitable for network layers with ReLU activation functions. These initialization methods can ensure that the distribution of parameters is reasonable and avoid the problem of gradient disappearance or gradient explosion during training.
[0150] In addition, in the early stages of model training, you can choose to freeze the weights of some networks (such as the shallow part of the backbone network) and only train the parameters of task-related modules (such as DPBlock, DFPN, DetBlock). This weight freezing strategy can reduce computational overhead and improve the learning efficiency of the model on task-related features.
[0151] 2. Optimizer and Learning Rate Strategy
[0152] The selection of optimizers and the design of learning rate strategies are important links in model training that cannot be ignored and have a direct impact on the convergence speed and performance of the model. In HardRockNet training, you can choose an optimizer suitable for the target detection task, such as AdamW or SGD, and adjust the corresponding hyperparameters according to the characteristics of the task.
[0153] AdamW is an improved Adam optimizer that effectively prevents overfitting by adding a weight decay term. For the AdamW optimizer, set the initial learning rate to 1e-4 or 5e-4 and the weight decay to 1e-2 or 1e-3. This optimizer is suitable for scenarios with limited hardware resources or small data volumes, and can provide faster convergence and better performance.
[0154] Another option is the SGD optimizer, which is a classic optimization method that performs stably in large-scale object detection tasks. For the SGD optimizer, the initial learning rate is set to 0.01 or 0.1, the momentum is set to 0.9, and the weight decay is set to 5e-4. Although SGD converges more slowly, it often outperforms AdamW in the final performance.
[0155] In order to dynamically adjust the learning rate, a learning rate scheduler can be used. Common strategies include cosine annealing scheduler and piecewise descent scheduler. The cosine annealing scheduler will gradually decay the learning rate according to the cosine curve, effectively preventing the instability caused by too high a learning rate in the later stages of training. The piecewise descent scheduler reduces the learning rate by a certain proportion after a fixed number of training rounds (for example, reducing the learning rate to 0.1 times the original value every 10 or 20 epochs). It is simple to operate and suitable for stable training tasks.
[0156] In addition, in the early stages of training, the Warm-up strategy can be used to gradually increase the learning rate, thereby avoiding the gradient explosion problem caused by an excessively large learning rate in the initial stage. The Warm-up strategy is particularly suitable for situations where the model weights are poorly initialized or the scale of training data is small.
[0157] Through the above detailed model initialization steps, including loading pre-trained weights, module weight initialization, optimizer selection and learning rate scheduler design, the training efficiency and performance of the HardRockNet model can be effectively improved. This initialization method can not only provide a good foundation for small target detection tasks, but also ensure the robustness and detection accuracy of the model in complex backgrounds and overlapping target scenes.
[0158] 3. Training process
[0159] Objective function definition:
[0160] In order to solve the problems of small target detection difficulty, complex background and target overlap in rock detection, the objective function adopts the multi-task joint loss method, combining classification loss, bounding box regression loss and background separation optimization, and is defined as follows:
[0161] ;
[0162] in, , , It is the weight to balance the importance of each loss term.
[0163] 1. Classification Loss
[0164] The classification loss is used to evaluate the accuracy of target category prediction. The improved Focal Loss is used to solve the problem of small target sample imbalance. The formula is as follows:
[0165]
[0166] is the predicted probability of the target category; is the category balancing factor; It is a parameter that controls the weights of easy-to-classify samples and hard-to-classify samples. Through the weighted mechanism of Focal Loss, the model's ability to classify small and overlapping objects can be improved.
[0167] 2. Bounding Box Regression Loss
[0168] The bounding box regression loss is used to optimize the difference between the predicted bounding box and the true bounding box. CIoU Loss (Complete IoU Loss) is used to more comprehensively describe the difference between overlapping targets. The formula is as follows:
[0169] ;
[0170] It is the intersection of the predicted box and the true box; Represents the Euclidean distance between the center points of the predicted box and the true box; is the diagonal length of the bounding box; is a shape consistency measure, Is a balance parameter. CIoU Loss not only considers IoU, but also introduces the consistency optimization of center point distance and aspect ratio, which is more suitable for scenes with complex background and overlapping targets.
[0171] 3. Background separation optimization
[0172] Background separation optimization aims to reduce the interference of complex background on target detection, using background suppression loss, which is defined as:
[0173] ;
[0174] in: The true label represents background (0) or target (1); is the probability predicted by the model; is the total number of samples.
[0175] Through this loss function, the model's prediction of the background area is explicitly optimized to avoid misjudgment between the background area and small targets.
[0176] Combining the above three parts, the final objective function can be expressed as:
[0177] ;
[0178] , , It can be adjusted according to the experimental results. The initial setting is generally This objective function can effectively improve the accuracy of stone target detection, especially the adaptability to small targets and complex background scenes, by jointly optimizing classification, bounding box regression and background separation.
[0179] Training steps:
[0180] In the process of model training, reasonable training steps are crucial, which directly affects the convergence speed and final performance of the model. This training adopts a phased training strategy. The following is a detailed description of the training steps:
[0181] 1) Rapid warm-up in the early stage (Warm-up)
[0182] In the early stages of training, we choose to use a larger learning rate to quickly warm up the model. The purpose of this stage is to accelerate the initial convergence process of the model so that the model can quickly capture the rough features in the data at the beginning. In this process, a larger learning rate helps to skip the local optimal solution and accelerate the update of weights. However, too large a learning rate may cause instability in training, so after the warm-up phase, the learning rate will gradually decrease to avoid oscillations during training.
[0183] 2) Entering the stable training phase
[0184] After the initial rapid warm-up, the training will enter a stable phase. At this time, the learning rate is gradually reduced to an appropriate range (for example, through a learning rate decay strategy or using a scheduler) to ensure that the model can be carefully optimized and achieve better generalization performance in the later stages of training. In this phase, the change in learning rate is usually adjusted according to certain rules (such as exponential decay, step decay, etc.) to avoid jumping out of the optimal solution too quickly in the later stages of training.
[0185] 3) Hardware configuration and batch size
[0186] This training used the NVIDIA 4090 GPU, which provides powerful computing power and video memory resources. To fully utilize the computing power of the GPU and improve training efficiency, we set a larger batch size (Batch Size) of 64. A larger batch size can improve the effect of data parallelism and reduce training time, while increasing the demand for memory and video memory during training. In order to ensure efficient training within the scope of hardware resources, the choice of Batch Size is determined based on the video memory capacity and computing power of the 4090 GPU.
[0187] 4) Training Cycles (Epochs)
[0188] This training is set to 200 epochs, that is, the model will traverse the entire training data set 200 times. During the training process of each epoch, the model will continuously update parameters based on the feedback of the loss function. The selection of the training cycle takes into account the size of the data set and the speed of model convergence. Through a longer training cycle, the model can gradually approach the global optimal solution, reduce the risk of overfitting, and improve the performance of the model on the validation set and test set.
[0189] 5) Mixed Precision Training
[0190] To speed up the training process and improve the utilization of hardware resources, we use mixed precision training technology. Mixed precision training uses 16-bit floating precision (FP16) instead of traditional 32-bit floating precision (FP32), thereby reducing video memory usage and accelerating calculation speed. Through mixed precision training, the NVIDIA 4090 GPU can improve training efficiency without significantly reducing accuracy. In addition, mixed precision training also reduces bandwidth consumption during calculations, further improving training performance.
[0191] 6) Summary and objectives
[0192] In general, the training steps range from rapid warm-up in the early stage, to fine-tuning in the stable stage, to strategies for efficient use of hardware resources (such as larger batch sizes and mixed precision training), aiming to accelerate the training process, improve the convergence speed of the model, and ultimately obtain a model with good generalization capabilities. Through reasonable learning rate scheduling, hardware optimization, and precision control, the goal of this training is to shorten the training time as much as possible while improving the effect of the model in real application scenarios.
[0193] Verification and tuning:
[0194] During the training process, it is crucial to ensure that the performance of the model continues to improve and reach the desired accuracy. Assuming that our goal is to increase the accuracy of the model to above 0.95, the following are the verification and tuning steps for this goal, including specific data and strategies:
[0195] 1) Validation set evaluation
[0196] After each training cycle (Epoch), we will use the validation set to evaluate the performance of the model. In our training process, assume that the evaluation results of the model at the 50th Epoch are as follows:
[0197] Average precision (mAP): At the 50th Epoch, the mAP of the model is 0.92, indicating that the overall accuracy of the model is already high, especially in the multi-category object detection task.
[0198] Recall: The recall rate is 0.87, which means that the model can successfully identify 87% of the real targets. A higher recall rate helps reduce missed detections and improves the coverage of target detection.
[0199] Precision: The accuracy reached 0.93 at the 50th Epoch, indicating that 93% of the samples predicted by the model as targets were correct, with fewer false positives and a higher reliability of the model.
[0200] But in order to further improve the accuracy and reach above 0.95, we will take the following verification and tuning steps.
[0201] Tuning hyperparameters to improve accuracy
[0202] In order to improve the accuracy, we need to carefully adjust the model's hyperparameters. Assuming the goal is to increase the accuracy to above 0.95, the specific adjustment strategy is as follows:
[0203] Learning Rate: During training, if the accuracy of the model stagnates or grows slowly, you can try to gradually adjust the learning rate. For example, if the current learning rate is 0.001, you can try to reduce it to 0.0005, or use a learning rate scheduler (such as ReduceLROnPlateau) to automatically reduce the learning rate when the performance on the validation set no longer improves. This can help the model to be carefully optimized in the later stages of training and avoid missing the optimal solution.
[0204] Weight Balancing Factors: For cases of class imbalance, if the number of samples in some classes is small, the class weights in the loss function can be adjusted appropriately. For example, the weights for small object classes can be increased to ensure that the model's recognition ability for these classes does not decrease. Suppose we set the weight of the small object class to 2.0 and the weight of the large object class to 1.0 to improve the detection accuracy of these small objects that are difficult to identify.
[0205] Batch Size: Batch size has an important impact on training speed and model accuracy. Larger batch sizes generally provide more stable gradient estimates, but require more video memory. Assuming the current batch size is 64, if there is enough video memory, the batch size can be increased to 128 to further improve training efficiency and stability. When increasing the batch size, it is usually necessary to adjust the learning rate appropriately.
[0206] 2) Prevent overfitting
[0207] To prevent the model from overfitting when improving accuracy and to ensure its generalization ability on the validation set and test set, we took the following measures:
[0208] Early Stopping: By setting up an early stopping mechanism, we ensure that training is stopped in time when the model performance no longer improves to prevent the model from overfitting. For example, when the mAP on the validation set does not improve significantly within 5 consecutive epochs, we will stop training to avoid wasting computing resources and reduce the risk of overfitting.
[0209] Regularization: To further enhance the generalization ability of the model, we used L2 regularization and Dropout. Assuming that the coefficient of L2 regularization is set to 0.0001 and the Dropout rate is 0.5, it means that 50% of the neuron connections are randomly discarded during each training. These regularization techniques help prevent the model from over-relying on the detailed features in the training set, thereby enhancing the generalization ability of the model.
[0210] Data Augmentation: To enhance the robustness of the model, we used data augmentation techniques such as random cropping, rotation, flipping, brightness adjustment, etc. This not only increases the diversity of the training data, but also enables the model to be trained under different conditions, improving its performance on unseen data. For example, we used a 50% flip probability, a rotation angle range of -30° to 30°, and a crop ratio between 0.8 and 1.0.
[0211] 3) Monitoring and adjustment
[0212] Assume that after adjustment, at the 100th Epoch, the evaluation results of the model are as follows:
[0213] Average Precision (mAP): The mAP of the model has been improved to 0.94, indicating that the performance of the model in multi-category detection tasks has been significantly improved.
[0214] Recall rate: The recall rate is increased to 0.90, the model's recognition rate of the target is increased, and the number of missed detections is reduced.
[0215] Accuracy: The accuracy has reached 0.95, which is the expected goal. This means that the model's prediction of positive samples is very accurate and the false positive rate is very low. At this point, we can judge that the performance of the model has reached or is close to the optimal state and stop training.
[0216] Through the above adjustment strategy and verification process, we successfully improved the accuracy of the model to above 0.95, ensuring that the model has higher precision and stronger generalization ability in detection tasks. This process not only improves the performance of the model on the training set, but also enables the model to have good effects in practical applications, especially when dealing with complex backgrounds and small targets, and can obtain more accurate detection results.
[0217] In summary, the grinding rock target detection method based on the HardRockNet network structure proposed in the present invention provides an efficient, accurate and real-time solution to the problems of small rock targets, complex background, target overlap and poor real-time performance in the prior art. The method has the following significant effects and characteristics:
[0218] 1. Excellent small target detection capability:
[0219] The present invention adopts the HardRockNet network structure and specially designs the DPBlock module, which effectively improves the detection capability of small targets through multi-scale feature extraction and deep convolution operation. The module can capture the local fine-grained features of small targets, thereby significantly improving the detection capability of small changes of hard rock targets in the ore background.
[0220] In the complex ore background, even if the stubborn rock target is small, HardRockNet can still achieve high detection accuracy and reduce missed detections.
[0221] 2. Powerful background interference suppression capability:
[0222] Through the DFPN (Dual Feature Pyramid Network) module, HardRockNet can effectively fuse multi-scale features and interactively enhance contextual information. This design effectively suppresses the interference of background noise on target detection, especially in complex grinding scenarios, and can effectively reduce the problem of false detection caused by background complexity.
[0223] Especially in the overlapping areas of ores and other minerals, DFPN enhances the model's perception of the target and improves its robustness to complex backgrounds.
[0224] 3. Dealing with target overlap:
[0225] The present invention uses the DetBlock module, adopts the method of step-by-step feature fusion and spatial information enhancement, and particularly effectively solves the problem of target overlap in the grinding process. DetBlock can refine the feature expression of the target, so that the network can distinguish overlapping targets and avoid the common misjudgment phenomenon in traditional methods.
[0226] This feature is particularly suitable for the overlapping detection of a large number of crushed stones and stubborn stones in the grinding process, ensuring the accurate identification and positioning of overlapping targets.
[0227] 4. Efficient real-time detection and feedback capabilities:
[0228] HardRockNet is designed with high computational efficiency and can quickly process image data and make predictions in a grinding environment with high real-time requirements. The network greatly improves the speed of target detection through efficient feature extraction and multi-scale fusion, providing strong support for real-time adjustment of the grinding process.
[0229] Using this method, the system can detect the distribution of stubborn stones in real time during the grinding process and adjust the process parameters according to the detection results, thereby improving the processing efficiency of the ore and reducing the grinding cost.
[0230] 5. Multi-task joint optimization to improve detection accuracy:
[0231] The present invention adopts a multi-task joint loss function (MultiTask Loss) to optimize the classification loss, bounding box regression loss and background suppression loss, so that the network can not only improve the classification accuracy of the target, but also accurately regress the bounding box position of the target, while reducing the confusion between complex background and small targets, further improving the detection accuracy.
[0232] Especially for small targets and overlapping targets in complex scenes, this loss function can help the network better locate and distinguish different targets, thereby improving the comprehensiveness and accuracy of detection.
[0233] 6. Special optimization for grinding process:
[0234] The detection method of the present invention is specially customized for the field of grinding, taking into account the particularity of ore and stubborn rock in the grinding process. Compared with the traditional general target detection method, HardRockNet can better cope with the unique challenges of the grinding process such as the wide variety of ores and complex backgrounds, and has stronger pertinence and adaptability.
[0235] In summary, the grinding rock target detection method based on the HardRockNet network structure proposed in the present invention has significantly improved the detection accuracy and efficiency of rock in the grinding process due to its strong small target detection capability, good background interference suppression capability, accurate target overlap processing, and superior real-time performance, providing strong support for the optimization and adjustment of the grinding process.
[0236] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A grinding rock target detection method based on HardRockNet network structure, characterized in that: The following steps are involved: Construct a target detection model for grinding stones, the target detection model includes an input layer, a multi-scale feature extraction layer, a DFPN feature fusion layer and an output layer connected in series in sequence; the multi-scale feature extraction layer includes a plurality of feature extraction sublayers connected in series in sequence and having different feature extraction scales, the output ends of the plurality of feature extraction sublayers are all connected to the input end of the DFPN feature fusion layer, and the plurality of feature extraction sublayers all use the DPBlock module for feature extraction; Collecting grinding pictures containing stubborn stones in different environments, and marking the stubborn stones on the grinding pictures to construct a training set, using the training set to train the target detection model, and using the trained target detection model to detect the stubborn stones on the grinding pictures; The DPBlock module includes: a global feature extraction unit, a local feature extraction unit and a fusion unit, and the output ends of the global feature extraction unit and the local feature extraction unit are connected to the input end of the fusion unit; The global feature extraction unit comprises: A first global feature extraction network, a first segmentation network, a first local feature extraction network, a first point-by-point convolution network, a first feature fusion network, a second point-by-point convolution network, and a first feature splicing network; the output end of the first global feature extraction network is connected to the input end of the first segmentation network, the output end of the first segmentation network is respectively connected to the input ends of the first local feature extraction network, the first point-by-point convolution network, and the first feature splicing network, the output end of the first local feature extraction network and the output end of the first point-by-point convolution network are both connected to the input end of the first feature fusion network, the output end of the first feature fusion network is connected to the input end of the first feature splicing network through the second point-by-point convolution network, and the output end of the first feature splicing network is connected to the input end of the fusion unit; The local feature extraction unit comprises: A second local feature extraction network, a second segmentation network, a second global feature extraction network, a third point-by-point convolutional network, a second feature fusion network, a fourth point-by-point convolutional network, and a second feature splicing network; the output end of the second local feature extraction network is connected to the input end of the second segmentation network, the output end of the second segmentation network is respectively connected to the input ends of the second global feature extraction network, the third point-by-point convolutional network, and the second feature splicing network, the output end of the second global feature extraction network and the output end of the third point-by-point convolutional network are both connected to the input end of the second feature fusion network, the output end of the second feature fusion network is connected to the input end of the second feature splicing network through the fourth point-by-point convolutional network, and the output end of the second feature splicing network is connected to the input end of the fusion unit; The fusion unit includes a third feature fusion network and a fifth point-by-point convolutional network; the third feature fusion network is connected to the output ends of the global feature extraction unit and the local feature extraction unit respectively; the output end of the third feature fusion network is connected to the input end of the fifth point-by-point convolutional network, and the output end of the fifth point-by-point convolutional network is connected to the input end of the DFPN feature fusion layer; The DFPN feature fusion layer includes: multiple feature pyramid branches, each of which corresponds to the multiple feature extraction sub-layers one by one, and each pyramid branch is used to extract the first feature of the output feature of the corresponding feature extraction sub-layer, and fuse and splice the first feature extracted by the pyramid branch with the first features extracted by other pyramid branches to output the corresponding fused feature.
2. The grinding rock target detection method based on the HardRockNet network structure according to claim 1 is characterized in that: The feature extraction sublayer includes a downsampling module and a DPBlock module connected in series, and the input layer is connected to the downsampling module at the head end of the multi-scale feature extraction layer; the output ends of the DPBlock modules are connected to the input ends of the DFPN feature fusion layer; and / or The multi-scale feature extraction layer includes a first feature extraction sublayer and a second feature extraction sublayer, the first feature extraction sublayer includes a first downsampling module and a first DPBlock module, the second feature extraction sublayer includes a second downsampling module and a second DPBlock module, the input end of the first downsampling module is connected to the output end of the input layer, the output end of the first downsampling module is connected to the input end of the first DPBlock module, the output end of the first DPBlock module is connected to the input end of the second downsampling module, the output end of the second downsampling module is connected to the input end of the second DPBlock module, and the output end of the first DPBlock module and the output end of the second DPBlock module are both connected to the input end of the DFPN feature fusion layer.
3. The grinding rock target detection method based on HardRockNet network structure according to claim 1 or 2, characterized in that: The DFPN feature fusion layer includes a first feature pyramid branch and a second feature pyramid branch; The first feature pyramid branch includes a third DPBlock module, a fifth point-by-point convolutional network, a fourth feature fusion network, a fourth DPBlock module, a sixth point-by-point convolutional network and a third feature splicing network, the output end of the third DPBlock module is connected to the input end of the fourth feature fusion network through the fifth point-by-point convolutional network, the output end of the third DPBlock module is also directly connected to the input end of the fourth feature fusion network, the output end of the fourth feature fusion network is connected to the input end of the fourth DPBlock module, the output end of the fourth DPBlock module is connected to the input end of the third feature splicing network through the sixth point-by-point convolutional network, and the output end of the fourth DPBlock module is also directly connected to the input end of the third feature splicing network; The second feature pyramid branch includes a fifth DPBlock module, a seventh point-by-point convolutional network, a fourth feature splicing network, a sixth DPBlock module, an eighth point-by-point convolutional network and a fifth feature fusion network, the output end of the fifth DPBlock module is connected to the input end of the fourth feature splicing network through the seventh point-by-point convolutional network, the output end of the fifth DPBlock module is also directly connected to the input end of the fourth feature splicing network, the output end of the fourth feature splicing network is connected to the input end of the sixth DPBlock module, the output end of the sixth DPBlock module is connected to the input end of the fifth feature fusion network through the eighth point-by-point convolutional network, and the output end of the sixth DPBlock module is also directly connected to the input end of the fifth feature fusion network; The third DPBlock module is jump-connected to the fourth feature splicing network, the fifth DPBlock module is jump-connected to the fourth feature fusion network, the fourth DPBlock module is jump-connected to the fifth feature fusion network, and the sixth DPBlock module is jump-connected to the third feature splicing network.
4. The grinding rock target detection method based on the HardRockNet network structure according to claim 3 is characterized in that: The output layer includes: a plurality of DetBlock modules, the plurality of DetBlock modules correspond one-to-one to the plurality of feature pyramid branches, and each DetBlock module performs stubborn stone target detection based on the fusion features output by the corresponding feature pyramid branch; The DetBlock module includes: a first channel adjustment convolution layer, a seventh DPBlock module, a fifth feature splicing network, a second channel adjustment convolution layer, an eighth DPBlock module, a sixth feature splicing network and a third channel adjustment convolution layer; The output end of the first channel adjustment convolution layer is connected to the input end of the fifth feature splicing network through the seventh DPBlock module, the output end of the first channel adjustment convolution layer is also directly connected to the input end of the fifth feature splicing network, the output end of the fifth feature splicing network is connected to the input end of the second channel adjustment convolution layer, the output end of the second channel adjustment convolution layer is connected to the input end of the sixth feature splicing network through the eighth DPBlock module, the output end of the second channel adjustment convolution layer is also directly connected to the input end of the sixth feature splicing network, and the output end of the sixth feature splicing network is connected to the input end of the third channel adjustment convolution layer.
5. The grinding rock target detection method based on HardRockNet network structure according to claim 4 is characterized in that: The loss function of the target detection model is satisfied by: ; ; ; ; in, is the classification loss, is the bounding box regression loss, Optimize loss for background separation, , , is the importance weight to balance each loss term, is the predicted probability of the target category; is the category balancing factor; It is a parameter that controls the weights of easy-to-classify samples and hard-to-classify samples; It is the intersection of the predicted box and the true box; Represents the Euclidean distance between the center points of the predicted box and the true box; Represents the coordinates of the center point of the prediction box, Indicates the coordinates of the center point of the real frame; is the diagonal length of the bounding box; is a shape consistency measure, is the equilibrium parameter; The true label representing the background or target; is the probability predicted by the model; is the total number of samples.
6. A computer system comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Traffic remote sensing target detection method
CN119169268A