Coal gangue image detection method, apparatus and device, and medium

Through the improved target detection network model, the HLAnet image enhancement module, MLKA attention module, etc. are used to extract and enhance the coal gangue image characteristics, which solves the problem of inaccurate detection when facing dense gangue, and achieves higher detection accuracy and processing efficiency.

CN120070974APending Publication Date: 2025-05-30SHENZHEN HIVT TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510137590.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When the target detection method based on Yolov9 faces highly dense gangue, there are adhesions and overlaps between coal gangues, resulting in inaccurate detection.

Method used

The improved object detection network model is adopted, including feature extraction network, which includes HLAnet image enhancement module, MLKA attention module, GI-a module, GI-b module, and GI-c module. Through these modules, image features are extracted and enhanced, the classification and positioning of coal gangue is performed.

Benefits of technology

Effectively distinguishing overlapping and obstructed coal from gangue improves the accuracy of detecting coal gangue, reduces the amount of parameters and floating point operations, and improves the processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070974A_ABST
    Figure CN120070974A_ABST
Patent Text Reader

Abstract

The invention relates to a coal gangue image detection method, device and equipment and a medium, and relates to the technical field of target detection.The method comprises the steps that coal gangue features of an input image are extracted through a target detection network model, a feature map is obtained, and the target detection network model comprises a feature extraction network; the feature extraction network comprises an HLAnet image enhancement module, an MLKA attention module, a G I-a module, a G I-b module and a G I-c module; the coal gangue of the feature map is classified and positioned, a detection result is output, and the detection result comprises the input image and the category and position information of the coal gangue. The overlapped and adhered coal and gangue are effectively distinguished through the improved target detection network model, fine features in the image can be better captured based on the scheme of deep learning, and the accuracy of coal gangue detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target detection, and particularly to a method, device, equipment and medium for detecting coal gangue images. Background Art

[0002] Currently, the target detection method based on Yolov9 has problems of adhesion and overlap between coal gangues when facing highly dense gangues, resulting in inaccurate detection. In coal gangue detection, there are large amounts of floating-point operations, slow inference speed, poor robustness and real-time performance. Relying solely on this method cannot solve the key problems, and a better detection scheme needs to be explored. Summary of the Invention

[0003] The present invention provides a method for detecting coal gangue images to solve the problem of inaccurate target detection caused by adhesion and overlap between coal gangues.

[0004] In a first aspect, the present invention provides a method for detecting coal gangue images, the method comprising:

[0005] Extracting coal gangue features of an input image through a target detection network model to obtain a feature map, wherein the target detection network model includes a feature extraction network, and the feature extraction network includes an HLAnet image enhancement module, an MLKA attention module, a GI-a module, a GI-b module, and a GI-c module;

[0006] Classifying and positioning the coal gangues in the feature map, and outputting a detection result, the detection result including the input image, the category and position information of the coal gangues.

[0007] In a second aspect, the present invention provides a device for detecting coal gangue images, comprising:

[0008] An extraction unit for extracting coal gangue features of an input image through a target detection network model to obtain a feature map, wherein the target detection network model includes a feature extraction network, and the feature extraction network includes an HLAnet image enhancement module, an MLKA attention module, a GI-a module, a GI-b module, and a GI-c module;

[0009] A classification unit for classifying and positioning the coal gangues in the feature map, and outputting a detection result, the detection result including the input image, the category and position information of the coal gangues.

[0010] In a third aspect, there is provided an electronic device, comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus;

[0011] The memory is used for storing a computer program;

[0012] A processor, when executing a program stored in a memory, implements the steps of the coal gangue image detection method according to any one of the embodiments of the first aspect.

[0013] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the coal gangue image detection method according to any one of the embodiments of the first aspect are implemented.

[0014] The above technical solutions provided by the embodiments of the present invention have the following advantages compared with the prior art:

[0015] The improved object detection network model can effectively distinguish overlapping and adhered coal and gangue, and the deep learning-based solution can better capture the subtle features in the image, improving the accuracy of coal gangue detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The drawings here are incorporated into the description and form a part of this description, showing the embodiments in line with the present invention, and are used together with the description to explain the principles of the present invention.

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0018] Figure 1 It is a schematic flowchart of a coal gangue image detection method provided by an embodiment of the present invention;

[0019] Figure 2 It is a schematic sub-flowchart of a coal gangue image detection method provided by an embodiment of the present invention;

[0020] Figure 3 It is a schematic sub-flowchart of a coal gangue image detection method provided by an embodiment of the present invention;

[0021] Figure 4 It is a schematic sub-flowchart of a coal gangue image detection method provided by an embodiment of the present invention;

[0022] Figure 5 It is a schematic sub-flowchart of a coal gangue image detection method provided by an embodiment of the present invention;

[0023] Figure 6 It is a schematic sub-flowchart of a coal gangue image detection method provided by an embodiment of the present invention;

[0024] Figure 7Schematic diagram of a sub - process of a coal gangue image detection method provided by an embodiment of the present invention;

[0025] Figure 8 Schematic diagram of a process of another coal gangue image detection method provided by an embodiment of the present invention;

[0026] Figure 9 Schematic diagram of the structure of a coal gangue image detection device provided by an embodiment of the present invention;

[0027] Figure 10 Schematic diagram of the structure of a computer device provided by an embodiment of the present invention;

[0028] Figure 11 An input image provided by an embodiment of the present invention;

[0029] Figure 12 is Figure 11 The effect diagram after being processed by the down - sampling module;

[0030] Figure 13 is Figure 12 The effect diagram after being processed by the MLKA attention module;

[0031] Figure 14 An input image provided by an embodiment of the present invention;

[0032] Figure 15 is Figure 14 The effect diagram after being processed by the down - sampling module;

[0033] Figure 16 is Figure 15 The effect diagram after being processed by the MLKA attention module;

[0034] Figure 17 An input image provided by an embodiment of the present invention;

[0035] Figure 18 is Figure 17 The effect diagram after being processed by the down - sampling module;

[0036] Figure 19 is Figure 18 The effect diagram after being processed by the MLKA attention module. Detailed implementation manners

[0037] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0038] Example 1

[0039] Figure 1 The present invention provides a flow chart of a method for detecting coal gangue images. The present invention provides a method for detecting coal gangue images. Figure 1 The gangue image detection method includes the following steps S101-S102.

[0040] S101, extracting gangue features of an input image through a target detection network model to obtain a feature map, wherein the target detection network model includes a feature extraction network, and the feature extraction network includes an HLAnet image enhancement module, an MLKA attention module, a GI-a module, a GI-b module, and a GI-c module.

[0041] In the specific implementation, the target detection network model combines Ghost-inceptionV2 convolution, MLKA attention module and HLAnet image enhancement network to establish an improved GI-yolov9 target detection network. Through this network, not only can the detection and identification of coal and gangue be realized, but also the detection and identification of coal and gangue can be realized with high accuracy, which greatly reduces the number of parameters and floating-point operations and improves processing efficiency.

[0042] HLAnet Image Enhancement: HLAnet Image Enhancement is an image preprocessing technique based on deep learning, which aims to improve the quality and details of images, especially in low contrast, noisy or blurred conditions. By using a specific network architecture, HLAnet can automatically learn useful information in images and enhance the clarity of images, making objects in images easier to identify. The network optimizes the multi-layer features of the image, restores details and reduces noise, and is particularly suitable for target detection or image analysis tasks in complex or non-ideal environments. HLAnet image enhancement plays an important role in improving the quality of visual information and improving the accuracy of subsequent detection algorithms. See Figures 11 - 19, MLKA Attention Mechanism: The MLKA (Multi-level Knowledge Attention) attention mechanism is an attention mechanism based on multi-level knowledge fusion, aiming to weight image features so that the network can focus more on key regions or important features. By introducing feature information at multiple levels, this mechanism enhances the network's perception of information at different levels, optimizing the expression and importance determination of features. Through the effective fusion of low-level features and high-level semantic information, the MLKA mechanism can improve the network's performance in complex environments. Especially in tasks such as object detection, it can effectively improve the detection accuracy, reduce the interference of irrelevant regions, and thus enhance the model's recognition ability for targets such as coal and gangue. The GI-a module is an attention module based on gradient information. It analyzes the gradient information of the input feature map and adaptively adjusts the weights of the feature map, thereby increasing the network's attention to important features. The GI-a module is used to enhance the network's feature extraction ability for target objects. Especially in complex backgrounds and multi-object scenarios, it can more accurately detect the position and category of target objects. The GI-b module is a module based on multi-scale feature fusion. It can fuse feature maps of different scales to obtain richer feature information. The GI-b module is used to improve the network's detection ability for target objects of different scales. Especially when dealing with small target objects, it can more effectively extract the feature information of small targets, thereby improving the detection accuracy. The GI-c module is a lightweight module based on convolutional neural networks. By optimizing the structure and parameters of the convolutional kernel, it reduces the computational amount and the number of parameters of the network, thereby improving the network's operation efficiency. The GI-c module is used to improve the network's operation speed on the premise of ensuring the detection accuracy. Especially on resource-constrained devices, it can perform object detection more quickly.

[0043] In one embodiment, refer to Figure 2 , Figure 2 is a schematic sub-process diagram of a coal and gangue image detection method provided by an embodiment of the present invention. The feature extraction network includes a convolutional layer, a first feature layer, a second feature layer, and a third feature layer. The above step S101 includes steps S201-S205:

[0044] In specific implementation, the convolutional layer includes Ghost-inceptionV2 convolution. Ghost convolution is a technique for optimizing convolutional operations, which improves the efficiency of the network by reducing the computational load. Traditional convolutional operations process the input feature map using standard convolutional kernels to generate high-dimensional feature maps, while Ghost convolution generates "ghost features" to replace some redundant computations. Specifically, Ghost convolution first extracts features using a small number of convolutional operations, and then generates additional "ghost" features through a specific generation strategy, thereby achieving feature expansion without increasing the computational burden. This method can effectively reduce the number of model parameters and computational volume, improve the operation efficiency of the network, and is particularly suitable for real-time detection and resource-constrained environments. The first feature layer, the second feature layer, and the third feature layer respectively refer to the feature maps extracted at different depths in the network. These feature maps have different resolutions and semantic information and are used for subsequent object detection and classification tasks.

[0045] S201, perform convolution on the input image through the convolutional layer to obtain a convolutional image.

[0046] In specific implementation, conv is the abbreviation of the convolutional layer and is used to extract image features in YOLOv9. The convolutional layer performs a convolutional operation by sliding a convolutional kernel over the image to extract the feature map of the image.

[0047] In this embodiment, the feature map after the convolutional layer performs convolution on the input image is denoted as the convolutional image.

[0048] S202, perform downsampling on the convolutional image through the first feature layer to obtain a first feature image.

[0049] In specific implementation, the first feature layer is a shallow feature layer with a relatively high resolution, which can capture detailed information in the image, such as edges and textures. The semantic information of the shallow feature layer is weak: due to the small receptive field, it mainly contains local feature information and is sensitive to the detection of small targets. Downsampling, also known as downsampling (ADown), is a commonly used technique in deep learning models, which helps the model capture the features of the image at a higher level while reducing the computational volume. The ADown module realizes downsampling through convolutional operations and stride adjustment. Its design focuses on lightweight and information retention, and can retain as much image information as possible while reducing the resolution of the feature map. In the network structure of YOLOv9, the ADown module can be integrated into the Backbone and Head, providing multiple configuration options to adapt to different improvement methods.

[0050] In this embodiment, perform downsampling on the convolutional image through the first feature layer to obtain a feature map with reduced resolution, denoted as the first feature image.

[0051] In one embodiment, see Figure 3 , Figure 3 A schematic diagram of a sub-process of a coal gangue image detection method provided by an embodiment of the present invention. The first feature layer includes an HLAnet image enhancement module, a downsampling module, and an MLKA attention module. The above step S202 includes steps S301-S303:

[0052] S301, preprocessing the convolution image through the HLAnet image enhancement module to obtain a first enhanced image.

[0053] In the specific implementation, preprocessing includes low contrast, noise interference or blur. HLAnet can automatically learn useful information in the image and enhance the clarity of the image by using a specific network architecture, making the target in the image easier to identify. The network optimizes the multi-layer features of the image, restores details and reduces noise, and is particularly suitable for target detection or image analysis tasks in complex or undesirable environments. The HLAnet image enhancement module plays an important role in improving the quality of visual information and improving the accuracy of subsequent detection algorithms.

[0054] In this embodiment, the clarity of the convolution image is improved by the HLAnet image enhancement module, so that the gangue in the image is easier to identify, and the image with enhanced clarity is recorded as the first enhanced image.

[0055] S302: Downsample the first enhanced image by the downsampling module to obtain a first feature point heat map.

[0056] In a specific implementation, the first enhanced image is downsampled by a downsampling module to obtain a feature map with reduced resolution, and the gangue identified in the feature map is highlighted to obtain a heat map marked with gangue, which is recorded as a first feature point heat map.

[0057] S303, enhancing the contour features of the gangue in the first feature point heat map through the MLKA attention module to obtain a first feature image.

[0058] In a specific implementation, the gangue information in the first feature point heat map is enhanced through the MLKA attention module, guiding the network to focus on the target area, such as the edges and details of coal blocks and gangue.

[0059] S203, downsampling the first feature image through a second feature layer to obtain a second feature image.

[0060] In specific implementation, the second feature layer is the middle feature layer, whose resolution is between that of the shallow feature layer and the deep feature layer. It can retain certain detailed information while having richer semantic information. The semantic information of the middle feature layer is strong: through multiple convolutional and pooling operations of the network, the middle feature layer can learn more abstract feature representations, which is more effective for the detection of medium-sized objects.

[0061] In this embodiment, the first feature image is downsampled by the second feature layer to obtain a feature map with reduced resolution, denoted as the second feature image.

[0062] In one embodiment, referring to Figure 4 , Figure 4 is a schematic diagram of a sub-process of a coal gangue image detection method provided by an embodiment of the present invention. The second feature layer includes a GI-a module, a downsampling module, and an MLKA attention module. The above step S203 includes steps S401-S403:

[0063] S401, enhance the position and category features of the first feature image through the GI-a module to obtain a second enhanced image.

[0064] In specific implementation, the position and category of coal gangue are detected through the GI-a module in complex background and multi-object scenarios.

[0065] S402, downsample the second enhanced image through the downsampling module to obtain a second feature point heat map.

[0066] In specific implementation, the second enhanced image is downsampled through the downsampling module to obtain a feature map with reduced resolution, and the identified coal gangue in the feature map is highlighted to obtain a heat map marked with coal gangue, denoted as the second feature point heat map.

[0067] S403, enhance the contour features of the coal gangue in the second feature point heat map through the MLKA attention module to obtain a second feature image.

[0068] In specific implementation, the coal gangue information in the second feature point heat map is enhanced through the MLKA attention module to guide the network to focus on the target area, such as the edges and details of coal blocks and coal gangue.

[0069] S204, downsample the second feature image through the third feature layer to obtain a third feature image.

[0070] In specific implementation, the third feature layer is the deep feature layer. After multiple downsampling operations, the resolution of the deep feature layer is low, but it has the strongest semantic information. The semantic information of the deep feature layer is rich: it can learn the high-level semantic features of the objects in the image, which plays an important role in the detection and classification of large objects.

[0071] In this embodiment, the second feature image is downsampled by the third feature layer to obtain a feature map with reduced resolution, denoted as the third feature image.

[0072] In one embodiment, refer to Figure 5 , Figure 5 which is a schematic diagram of a sub - process of a coal gangue image detection method provided by an embodiment of the present invention. The third feature layer includes a GI - b module, a downsampling module, an MLKA attention module, and a GI - c module. The above step S204 includes steps S501 - S503:

[0073] S501, enhance the features of small coal gangues in the second feature image through the GI - b module to obtain a third enhanced image.

[0074] In specific implementation, the GI - b module enhances the network's detection ability for coal gangues of different scales. Especially when dealing with small coal gangues, it can more effectively extract the feature information of small coal gangues, thereby improving the detection accuracy.

[0075] S502, downsample the third enhanced image through the downsampling module to obtain a third feature point heat map.

[0076] In specific implementation, the third enhanced image is downsampled through the downsampling module to obtain a feature map with reduced resolution, and the identified coal gangues in the feature map are highlighted to obtain a heat map marked with coal gangues, denoted as the third feature point heat map.

[0077] S503, enhance the contour features of small coal gangues in the third feature point heat map through the MLKA attention module and the GI - c module to obtain a third feature image.

[0078] In specific implementation, the MLKA attention module and the GI - c module enhance the small coal gangue information in the third feature point heat map, guiding the network to focus on the target area, such as the edges and details of coal blocks and coal gangues. Among them, there are more small coal gangues identified in the third feature point heat map. By optimizing the structure and parameters of the convolution kernel through the GI - c module, the calculation amount and the number of parameters of the network are reduced, thereby improving the operation efficiency of the network.

[0079] S205, use the first feature image, the second feature image, and the third feature image as the feature map.

[0080] In specific implementation, the first feature image, the second feature image, and the third feature image obtained in steps S202 - S204 are used as the feature map of step S101 and supplied to step S102 for further processing.

[0081] S102. Classify and locate the coal gangue in the feature map, and output the detection result, where the detection result includes the input image, the category and location information of the coal gangue.

[0082] In specific implementation, the coal gangue in the feature map is classified and located through the object detection network model, and the category and location information of the coal gangue are marked in the image.

[0083] In one embodiment, refer to Figure 6 , Figure 6 which is a schematic sub - process diagram of a coal gangue image detection method provided by an embodiment of the present invention. The object detection network model further includes a fusion network and a classification network. The above step S102 includes steps S601 - S602:

[0084] S601. Fusion the feature map through the fusion network to obtain a fused feature map.

[0085] In specific implementation, the fusion network refers to the Neck part of the object detection network model, which is used to further fuse and process the features extracted by the Backbone part to improve the feature expression ability.

[0086] In this embodiment, the fusion network further fuses and processes the features extracted in step S101 to improve the feature expression ability.

[0087] In one embodiment, refer to Figure 7 , Figure 7 which is a schematic sub - process diagram of a coal gangue image detection method provided by an embodiment of the present invention. The above step S601 includes steps S701 - S705:

[0088] S701. Fuse the features of the second feature image and the features of the third feature image to obtain a fused feature.

[0089] S702. Fuse the features of the first feature image and the fused feature to obtain a first fused feature map.

[0090] S703. Fuse the features of the second feature image, the fused feature and the features of the first fused feature map to obtain a second fused feature map.

[0091] S704. Fuse the features of the third feature image, the fused feature and the features of the second fused feature map to obtain a third fused feature map.

[0092] S705. Use the first fused feature map, the second fused feature map and the third fused feature map as the fused feature map.

[0093] In specific implementation, steps S701 - S705 respectively fuse the feature maps extracted in step S101. Refer to Figure 8 , Adown is the downsampling module in YOLOv9, which is used to reduce the spatial dimension of the feature map. Downsampling is a commonly used technique in deep learning models, which helps the model capture the features of the image at a higher level while reducing the computational cost. The ADown module achieves downsampling through convolutional operations and stride adjustment. Its design focuses on lightweight and information retention, and can retain as much image information as possible while reducing the resolution of the feature map. In the network structure of YOLOv9, the ADown module can be integrated into the Backbone and Head, providing multiple configuration options to adapt to different improvement methods. Upsample is the upsampling operation, which is used in YOLOv9 to restore the low - resolution feature map to a higher resolution. Upsampling is usually used to restore the size of the feature map in the network for subsequent fusion or prediction operations. In the network structure of YOLOv9, the Upsample operation may be used to fuse feature maps at different levels to improve the network's detection ability for targets of different scales. Concat is used to fuse the output feature maps of multiple convolutional layers, thereby achieving more efficient feature extraction and utilization. This feature fusion helps improve the model's representation ability and generalization ability, enabling the model to better adapt to different datasets and tasks.

[0094] S602, classify and locate the targets of the fused feature map through the classification network, and mark the category and location information of the targets in the input image. The input image after marking is used as the detection result.

[0095] In specific implementation, the classification network refers to the Head part of the target detection network model, which is used to perform final target detection and classification on the features after being processed by the Neck.

[0096] In this embodiment, the classification network performs final target detection and classification on the features processed in step S601, and marks the detected coal gangue in the image.

[0097] The embodiments of the present invention can achieve the following advantages:

[0098] The improved target detection network model can effectively distinguish overlapping and adhered coal and gangue, and the deep - learning - based solution can better capture the subtle features in the image, improving the accuracy of coal gangue detection.

[0099] Refer to Figure 9 The embodiments of the present invention also provide a coal gangue image detection device 400, which includes an extraction unit 401 and a classification unit 402.

[0100] An extraction unit 401 is configured to extract the gangue features of the input image through a target detection network model to obtain a feature map. The target detection network model includes a feature extraction network, and the feature extraction network includes an HLAnet image enhancement module, an MLKA attention module, a GI-a module, a GI-b module, and a GI-c module.

[0101] In one embodiment, the feature extraction network includes a convolutional layer, a first feature layer, a second feature layer, and a third feature layer. The process of extracting the gangue features of the input image through the target detection network model to obtain a feature map includes:

[0102] Performing convolution on the input image through the convolutional layer to obtain a convolutional image;

[0103] Performing downsampling on the convolutional image through the first feature layer to obtain a first feature image;

[0104] Performing downsampling on the first feature image through the second feature layer to obtain a second feature image;

[0105] Performing downsampling on the second feature image through the third feature layer to obtain a third feature image;

[0106] Using the first feature image, the second feature image, and the third feature image as the feature map.

[0107] In one embodiment, the first feature layer includes an HLAnet image enhancement module, a downsampling module, and an MLKA attention module. The process of performing downsampling on the convolutional image through the first feature layer to obtain a first feature image includes:

[0108] Performing preprocessing on the convolutional image through the HLAnet image enhancement module to obtain a first enhanced image;

[0109] Performing downsampling on the first enhanced image through the downsampling module to obtain a first feature point heat map;

[0110] Enhancing the contour features of the gangue in the first feature point heat map through the MLKA attention module to obtain a first feature image.

[0111] In one embodiment, the second feature layer includes a GI-a module, a downsampling module, and an MLKA attention module. The process of performing downsampling on the first feature image through the second feature layer to obtain a second feature image includes:

[0112] Enhancing the position and category features of the first feature image through the GI-a module to obtain a second enhanced image;

[0113] The second enhanced image is downsampled by the downsampling module to obtain a second feature point heat map;

[0114] The contour features of the coal gangue in the second feature point heat map are enhanced by the MLKA attention module to obtain a second feature image.

[0115] The classification unit 402 is used to classify and locate the coal gangue in the feature map and output a detection result, where the detection result includes the input image, the category and position information of the coal gangue.

[0116] In one embodiment, the third feature layer includes a GI-b module, a downsampling module, an MLKA attention module, and a GI-c module. The downsampling of the second feature image by the third feature layer to obtain a third feature image includes:

[0117] The features of small coal gangue in the second feature image are enhanced by the GI-b module to obtain a third enhanced image;

[0118] The third enhanced image is downsampled by the downsampling module to obtain a third feature point heat map;

[0119] The contour features of small coal gangue in the third feature point heat map are enhanced by the MLKA attention module and the GI-c module to obtain a third feature image.

[0120] In one embodiment, the object detection network model further includes a fusion network and a classification network. The classification and location of the object in the feature map and the output of the detection result include:

[0121] The feature map is fused by the fusion network to obtain a fused feature map;

[0122] The object in the fused feature map is classified and located by the classification network, and the category and position information of the object are marked in the input image. The input image after marking is used as the detection result.

[0123] In one embodiment, the fusion of the feature map by the fusion network to obtain a fused feature map includes:

[0124] The features of the second feature image and the features of the third feature image are fused to obtain a fused feature;

[0125] The features of the first feature image and the fused feature are fused to obtain a first fused feature map;

[0126] The features of the second feature image, the fused feature, and the features of the first fused feature map are fused to obtain a second fused feature map;

[0127] Fuse the features of the third feature image, the fusion features, and the features of the second fusion feature map to obtain a third fusion feature map;

[0128] Use the first fusion feature map, the second fusion feature map, and the third fusion feature map as the fusion feature map.

[0129] As Figure 10 shown, Figure 10 FIG. is a schematic block diagram of a computer device provided by an embodiment of the present application. The computer device 500 can be a terminal or a server. Among them, the terminal can be an electronic device with communication functions such as a smart phone, a tablet computer, a notebook computer, a desktop computer, a personal digital assistant, and a wearable device. The server can be an independent server or a server cluster composed of multiple servers.

[0130] The computer device 500 includes a processor 502, a memory, and a network interface 505 connected through a system bus 501. Among them, the memory can include a non-volatile storage medium 503 and an internal memory 504.

[0131] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. When the computer program 5032 is executed, the processor 502 can be caused to execute a coal gangue image detection method.

[0132] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.

[0133] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can be caused to execute a coal gangue image detection method.

[0134] The network interface 505 is used for network communication with other devices. Those skilled in the art can understand that the above structure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device 500 to which the solution of the present application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0135] It should be understood that in the embodiments of the present application, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0136] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a storage medium, and the storage medium is a computer-readable storage medium. The computer program is executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.

[0137] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program.

[0138] The storage medium is a physical, non-transitory storage medium, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc, etc., which are various physical storage media that can store program codes. The computer-readable storage medium may be non-volatile or volatile.

[0139] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.

[0140] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of each unit is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0141] The steps in the method embodiments of the present invention can be adjusted, combined, and deleted according to actual needs. The units in the device embodiments of the present invention can be combined, divided, and deleted according to actual needs. In addition, in each embodiment of the present invention, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0142] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention.

[0143] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0144] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, provided that these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these changes and modifications.

[0145] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or substitutions within the technical scope disclosed by the present invention, and these modifications or substitutions should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for detecting coal gangue images, characterized in that: The method comprises: Extracting gangue features of an input image through a target detection network model to obtain a feature map, wherein the target detection network model includes a feature extraction network, and the feature extraction network includes an HLAnet image enhancement module, an MLKA attention module, a GI-a module, a GI-b module, and a GI-c module; The gangue in the feature image is classified and located, and a detection result is output, wherein the detection result includes the input image, the category and location information of the gangue.

2. The method according to claim 1, characterized in that The feature extraction network includes a convolution layer, a first feature layer, a second feature layer and a third feature layer. The target detection network model is used to extract the gangue features of the input image to obtain a feature map, including: Convolving the input image through the convolution layer to obtain a convolved image; Downsampling the convolution image through a first feature layer to obtain a first feature image; Downsampling the first feature image through a second feature layer to obtain a second feature image; Downsampling the second feature image through a third feature layer to obtain a third feature image; The first feature image, the second feature image, and the third feature image are used as the feature map.

3. The method according to claim 2, characterized in that The first feature layer includes an HLAnet image enhancement module, a downsampling module, and an MLKA attention module. The convolution image is downsampled by the first feature layer to obtain a first feature image, including: Preprocessing the convolution image by the HLAnet image enhancement module to obtain a first enhanced image; Downsampling the first enhanced image by the downsampling module to obtain a first feature point heat map; The contour features of the coal gangue in the first feature point heat map are enhanced by the MLKA attention module to obtain a first feature image.

4. The method according to claim 2, characterized in that: The second feature layer includes a GI-a module, a downsampling module, and an MLKA attention module. The first feature image is downsampled by the second feature layer to obtain a second feature image, including: Enhance the position and category features of the first feature image by the GI-a module to obtain a second enhanced image; Downsampling the second enhanced image by the downsampling module to obtain a second feature point heat map; The contour features of the gangue in the second feature point heat map are enhanced by the MLKA attention module to obtain a second feature image.

5. The method according to claim 2, characterized in that: The third feature layer includes a GI-b module, a downsampling module, an MLKA attention module and a GI-c module. The second feature image is downsampled by the third feature layer to obtain a third feature image, including: Enhance the features of small coal gangue in the second feature image by the GI-b module to obtain a third enhanced image; Downsampling the third enhanced image by the downsampling module to obtain a third feature point heat map; The contour features of small and medium-sized coal gangue in the third feature point heat map are enhanced by the MLKA attention module and the GI-c module to obtain a third feature image.

6. The method according to claim 2, characterized in that The target detection network model also includes a fusion network and a classification network, and the target of the feature map is classified and located, and the detection result is output, including: The feature maps are fused through the fusion network to obtain a fused feature map; The target of the fused feature map is classified and located through the classification network, and the category and location information of the target are marked in the input image, and the marked input image is used as the detection result.

7. The method according to claim 6, characterized in that The step of fusing the feature maps through the fusion network to obtain a fused feature map includes: Fusing the features of the second feature image and the features of the third feature image to obtain a fused feature; Fusing the features of the first feature image and the fused features to obtain a first fused feature map; Fusing the features of the second feature image, the fusion features, and the features of the first fusion feature map to obtain a second fusion feature map; Fusing the features of the third feature image, the fusion features, and the features of the second fusion feature map to obtain a third fusion feature map; The first fused feature map, the second fused feature map and the third fused feature map are used as the fused feature map.

8. A coal gangue image detection device, characterized in that: include: An extraction unit, used for extracting gangue features of an input image through a target detection network model to obtain a feature map, wherein the target detection network model includes a feature extraction network, and the feature extraction network includes an HLAnet image enhancement module, an MLKA attention module, a GI-a module, a GI-b module, and a GI-c module; The classification unit is used to classify and locate the gangue in the feature image and output a detection result, wherein the detection result includes the input image and the category and location information of the gangue.

9. A computer device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, for implementing the steps of the method described in any one of claims 1 to 7 when executing a program stored in a memory.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.