Lightweight target detection method and system for coal mine underground small target detection

By combining the DSC-ResNet network model and the lightweight feature pyramid network model CA-LFPN, the problems of low detection accuracy and high computational cost in small target detection in coal mines are solved, and efficient small target detection is achieved.

CN120783027APending Publication Date: 2025-10-14WUXI CITY COLLEGE OF VOCATIONAL TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510922272.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

Existing DETR detectors have problems with memory access and inference delay in small target detection in coal mines, making it difficult to effectively identify small targets. In addition, traditional target detection methods cannot effectively deal with the correlation and redundancy between feature maps, resulting in low detection accuracy.

Method used

The DSC-ResNet network model is combined with the lightweight feature pyramid network model CA-LFPN. The image is enhanced through bilateral filtering to extract high-level features with rich semantic information. The lightweight feature pyramid network model is used to improve the detection accuracy while reducing the computational cost.

Benefits of technology

The detection accuracy of small targets in coal mines is improved, the computational cost is reduced, and the sensitivity and accuracy of detection are maintained.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120783027A_ABST
    Figure CN120783027A_ABST
Patent Text Reader

Abstract

The invention relates to a lightweight target detection method and a lightweight target detection system for small target detection in an underground coal mine. The method comprises the following steps: acquiring an acquired coal mine small-target low-illumination source image, and configuring bilateral filtering parameters to process the low-illumination source image to obtain a low-illumination target image; inputting the low-illumination target image into the DSC-ResNet network model, and respectively extracting a first image feature, a second image feature, a third image feature and a fourth image feature through convolution kernels of different sizes; inputting the fourth image features into a DSC-ResNet network model, and generating a feature map through inverse linear projection; and processing in a lightweight feature pyramid network model CA-LFPN to generate a prediction bounding box, and completing target prediction. The DSC-ResNet network model can ensure the detection precision under the condition of reducing the calculation cost, so that the detection precision of target detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of target detection, in particular to a lightweight target detection method and system for underground coal mine small target detection. BACKGROUND

[0002] Underground coal mine railway small target detection mainly involves detecting small targets such as bolts, spikes, obstacles, etc. on the railway track in a complex underground coal mine environment. Accurate underground coal mine railway small target detection is crucial for ensuring safety, improving efficiency, and reducing the safety risks of underground coal mine railway autonomous driving. In traditional target detection methods, the current mainstream DETR detector has major challenges in actual application due to the need for a large number of memory accesses and inference delays. In addition, the DETR model has difficulty in effectively learning and recognizing small targets because the small targets occupy a small proportion of the image and have less feature information, resulting in slow model convergence and insufficient small target detection performance. In target detection methods, feature extraction often uses multi-scale feature fusion methods, mainly fusing deep and shallow features to give shallow features strong semantic information. Two common multi-scale feature fusion methods are parallel multi-branch networks and serial skip connection structures. Among them, parallel multi-branch networks usually use different convolutions to extract features from the same feature, and then fuse the extracted features by concatenation. This idea is evident in the Inception module of GoogLeNet, which uses multiple convolutions to extract features from the same feature map and then combines them in the channel dimension. Unlike the parallel concatenation method, the serial skip connection structure usually performs multi-scale fusion on the output of different layers in the backbone network. The feature pyramid network (FPN) achieves multi-scale feature fusion by upsampling high-level features to a uniform scale and then adding them to the bottom-level features.

[0003] However, in traditional target detection methods, images are often acquired based on normal lighting environments, and the correlation and redundancy problems between feature maps cannot be effectively addressed during target detection, resulting in low target detection accuracy. SUMMARY

[0004] Therefore, in order to solve the above technical problems, a lightweight target detection method and system for underground coal mine small target detection are provided, which can improve the detection accuracy of target detection.

[0005] A lightweight target detection method for underground coal mine small target detection, the method comprising:

[0006] The acquired coal mine underground small target low-illumination source image is processed based on the mine environment parameters to obtain a low-illumination target image.

[0007] The low-illumination target image is input into a DSC-ResNet network model, and first, second, third and fourth image features are extracted through different size convolution kernels.

[0008] The fourth image feature is input into the average pooling layer of the DSC-ResNet network model for pooling, and normalized processing is performed by using a Softmax function to obtain processed linear one-dimensional data.

[0009] The linear one-dimensional data and the position encoding vector are pixel-level added to convert into a first image sequence, the first image sequence is input into an encoder for processing, and a feature map is generated through inverse linear projection.

[0010] The first, second, third and fourth image features and the feature map are input into a lightweight feature pyramid network model CA-LFPN for processing to obtain an output image.

[0011] The output image is linearized and pixel-level added to the linear one-dimensional data in the encoder, and input into a decoder, a predicted bounding box is generated based on the output data of the decoder, and target prediction is completed.

[0012] In one of the embodiments, the mine environment parameters are collected, and the bilateral filter parameters are configured based on the mine environment parameters, including:

[0013] The ambient brightness of the coal mine is collected by a light sensor, and the dust concentration of the coal mine is measured by a dust sensor.

[0014] The ambient brightness and dust concentration are used as mine environment parameters, and the bilateral filter parameters are configured according to the mine environment parameters.

[0015] In one of the embodiments, the low-illumination source image is processed based on the bilateral filter parameters to obtain a low-illumination target image, including:

[0016] The spatial neighborhood window of the low-illumination source image is determined.

[0017] The normalized weight of the pixel value in the spatial neighborhood window is obtained, and the spatial domain Gaussian kernel function and the pixel domain Gaussian kernel function are determined based on the spatial neighborhood window.

[0018] According to the normalized weight, the spatial domain Gaussian kernel function and the pixel domain Gaussian kernel function, a low-illumination target image is obtained by performing a bilateral filtering process on the low-illumination source image.

[0019] According to the filtered pixel value, the low-illumination target image is obtained.

[0020] In one embodiment, the low-illumination target image is input into a DSC-ResNet network model, and first, second, third and fourth image features are extracted by using different sizes of convolution kernels.

[0021] The low-illumination target image is input into the DSC-ResNet network model, and a preliminary feature is generated by performing a convolution pooling operation on the encoder in the DSC-ResNet network model.

[0022] According to the preliminary feature, the low-illumination target image is spatially down-sampled and the number of channels is increased by using different sizes of convolution kernels, and an intermediate feature is obtained.

[0023] The intermediate feature is input into a DSC Block module to obtain an image extraction feature.

[0024] According to the image extraction feature, first, second, third and fourth image features are obtained.

[0025] In one embodiment, according to the image extraction feature, first, second, third and fourth image features are obtained, including:

[0026] The image extraction feature is input into a convolution kernel of a first size after being spliced, and is added at a pixel level with the intermediate feature to obtain a first image feature.

[0027] The first image feature is input into a convolution kernel of a second size after being processed by an encoder to perform a convolution pooling operation, spatial down-sampling and increasing the number of channels, and then being processed by a DSC Block module to obtain a second image feature.

[0028] The second image feature is input into a convolution kernel of a third size after being processed by an encoder to perform a convolution pooling operation, spatial down-sampling and increasing the number of channels, and then being processed by a DSC Block module to obtain a third image feature.

[0029] The third image feature is input into a convolution kernel of a fourth size after being processed by an encoder to perform a convolution pooling operation, spatial down-sampling and increasing the number of channels, and then being processed by a DSC Block module to obtain a fourth image feature.

[0030] In one of the embodiments, the DSC Block module includes a depth separable convolution layer, a standard convolution layer, a batch normalization layer, and an activation layer.

[0031] In one of the embodiments, the first image feature, the second image feature, the third image feature, the fourth image feature, and the feature map are input into a lightweight feature pyramid network model CA-LFPN for processing, including:

[0032] The first image feature, the second image feature, the third image feature, the fourth image feature, and the feature map are input into the lightweight feature pyramid network model CA-LFPN, and five groups of features are obtained after processing by an attention mechanism module.

[0033] The five groups of features are subjected to convolution operation, and the fourth image feature is subjected to up-sampling processing. The standard fourth image feature after up-sampling processing is added to the third image feature after convolution operation at a pixel level to obtain a standard third image feature.

[0034] The standard third image feature is subjected to up-sampling processing and added to the second image feature after convolution operation at a pixel level to obtain a standard second image feature.

[0035] The standard second image feature is subjected to up-sampling processing and added to the first image feature after convolution operation at a pixel level to obtain a standard first image feature.

[0036] In one of the embodiments, the output image is subjected to linearization operation and added to the linear one-dimensional data in the encoder at a pixel level, and is input into a decoder. A prediction bounding box is generated based on output data of the decoder, including:

[0037] After the standard first image feature, the standard second image feature, the standard third image feature, and the standard fourth image feature are subjected to convolution operation, a second image sequence is generated by linear projection.

[0038] The first image sequence and the second image sequence are added at a pixel level, and the added image sequence is input into a decoder. A prediction bounding box is generated based on output data of the decoder.

[0039] A lightweight target detection system for small target detection in a coal mine underground, the system comprising:

[0040] An image acquisition and processing module is configured to acquire a low-illumination source image of a small target in a coal mine underground, acquire mine environment parameters, configure bilateral filtering parameters based on the mine environment parameters, and process the low-illumination source image based on the bilateral filtering parameters to obtain a low-illumination target image.

[0041] The feature extraction module is configured to input the low-illumination target image into a DSC-ResNet network model, and extract a first image feature, a second image feature, a third image feature and a fourth image feature through different sizes of convolution kernels respectively.

[0042] The linear processing module is configured to input the fourth image feature into an average pooling layer of the DSC-ResNet network model for pooling, and perform normalization processing on the fourth image feature by using a Softmax function to obtain processed linear one-dimensional data.

[0043] The encoding module is configured to add the linear one-dimensional data and a position encoding vector at a pixel level to convert the linear one-dimensional data into a first image sequence, input the first image sequence into an encoder for processing, and generate a feature map through inverse linear projection.

[0044] The feature processing module is configured to input the first image feature, the second image feature, the third image feature, the fourth image feature and the feature map into a lightweight feature pyramid network model CA-LFPN for processing to obtain an output image.

[0045] The prediction module is configured to perform linearization operation on the output image, add the output image and the linear one-dimensional data in the encoder at a pixel level, input the output image into a decoder, generate a predicted bounding box based on output data of the decoder, and complete target prediction.

[0046] The lightweight target detection method and system for small target detection in a coal mine underground can realize image enhancement by performing bilateral filtering processing on the collected low-illumination source image. The DSC-ResNet network model can extract high-level features with rich semantic information. The lightweight feature pyramid network model can realize network lightweight and feature combination while minimizing information loss. The DSC-ResNet network model can ensure detection accuracy while reducing calculation cost, thereby improving the detection accuracy of target detection. BRIEF DESCRIPTION OF DRAWINGS

[0047] Figure 1 It is an application environment diagram of the lightweight target detection method for small target detection in a coal mine underground in an embodiment;

[0048] Figure 2 It is a flowchart of the lightweight target detection method for small target detection in a coal mine underground in an embodiment;

[0049] Figure 3 It is a structure diagram of the DSC-ResNet network model in an embodiment;

[0050] Figure 4A structural schematic diagram of a CA-LFPN network model in an embodiment;

[0051] Figure 5 A structural block diagram of a lightweight target detection system for underground coal mine small target detection in an embodiment;

[0052] Figure 6 A structural schematic diagram of an MFFL-DETR network model in an embodiment;

[0053] Figure 7 An internal structure diagram of a computer device in an embodiment. DETAILED DESCRIPTION

[0054] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.

[0055] It can be understood that the terms "first", "second", and the like used in the present application can be used herein to describe image features, image sequences, but these image features, image sequences are not limited by these terms. These terms are only used to distinguish the first image feature, image sequence from another image feature, image sequence. For example, without departing from the scope of the present application, the first image feature can be referred to as the second image feature, and similarly, the second image feature can be referred to as the first image feature. The first image feature and the second image feature are both image features, but they are not the same image feature.

[0056] The lightweight target detection method for underground coal mine small target detection provided by the embodiments of the present application can be applied in an application environment as shown in Figure 1 as shown in Figure 1As shown, the application environment includes a computer device 110. The computer device 110 can acquire the acquired low-illumination source image of the small target in the coal mine underground, acquire the mine environment parameter, configure the bilateral filtering parameter based on the mine environment parameter, and process the low-illumination source image based on the bilateral filtering parameter to obtain a low-illumination target image; the computer device 110 can input the low-illumination target image into the DSC-ResNet network model, and extract first image features, second image features, third image features and fourth image features through convolution kernels of different sizes; the computer device 110 can input the fourth image features into the average pooling layer of the DSC-ResNet network model for pooling, and perform normalization processing by using a Softmax function to obtain processed linear one-dimensional data; the computer device 110 can add the linear one-dimensional data and the position encoding vector at the pixel level to convert the linear one-dimensional data into a first image sequence, input the first image sequence into an encoder for processing, and generate a feature map through inverse linear projection; the computer device 110 can input the first image features, the second image features, the third image features, the fourth image features and the feature map into the lightweight feature pyramid network model CA-LFPN for processing to obtain an output image; the computer device 110 can perform linearization operation on the output image, add the linear one-dimensional data in the encoder at the pixel level, input into a decoder, generate a predicted bounding box based on the output data of the decoder, and complete target prediction. The computer device 110 can be, but is not limited to, various personal computers, notebook computers, smart phones, robots, unmanned aerial vehicles, tablet computers and the like.

[0057] In one embodiment, as shown in Figure 2 A lightweight target detection method for small target detection in coal mine underground is provided, comprising the following steps:

[0058] In step 202, the acquired low-illumination source image of the small target in the coal mine underground is acquired, the mine environment parameter is acquired, the bilateral filtering parameter is configured based on the mine environment parameter, and the low-illumination source image is processed based on the bilateral filtering parameter to obtain a low-illumination target image.

[0059] The computer device can acquire the low-illumination source image of the small target in the coal mine underground, specifically, an imaging device such as a camera can be used to first shoot the low-illumination source image, and then the low-illumination source image is preprocessed to generate a low-illumination target image after preprocessing. At the same time, in order to acquire the mine environment parameter, the surrounding environment information can be acquired when shooting the image.

[0060] In one embodiment, a lightweight target detection method for detecting small targets underground in coal mines is provided, which may also include a process of collecting environmental parameters and configuring bilateral filtering parameters. The specific process includes: collecting the ambient brightness underground in the coal mine through a light sensor, and measuring the dust concentration underground in the coal mine through a dust sensor; using the ambient brightness and dust concentration as mine environmental parameters, and configuring the bilateral filtering parameters according to the mine environmental parameters.

[0061] When capturing an image, the mine environment parameters corresponding to the low-light source image can be obtained, specifically, the current mine environment brightness and mine dust concentration can be obtained. In this embodiment, the current coal mine environment brightness (cd / m²) can be measured using a BH1750 light sensor, and the current coal mine environment dust concentration (mg / m3) can be measured using a GP2Y1014AU dust sensor. In addition, other methods can be used to obtain the coal mine environment brightness and the corresponding mine dust concentration.

[0062] Next, the computer device can configure bilateral filtering parameters based on the mine environment parameters. Bilateral filtering can then be performed on the low-light image based on the bilateral filtering parameters. During the bilateral filtering process, the spatial proximity and pixel value similarity of the low-light source image can be considered to smooth the low-light source image while preserving its edge information. In this embodiment, a low-light target image can be generated after bilateral filtering. When bilateral filtering is no longer required for the low-light source image, the low-light source image can be configured as the low-light target image.

[0063] In one embodiment, a lightweight target detection method for detecting small targets underground in coal mines is provided, which may also include a process of performing bilateral filtering processing. The specific process includes: determining a spatial neighborhood window of a low-illumination source image; obtaining the normalized weights of the pixel values ​​in the spatial neighborhood window, and determining a spatial domain Gaussian kernel function and a pixel domain Gaussian kernel function based on the spatial neighborhood window; performing bilateral filtering processing on the low-illumination source image according to the normalized weights, the spatial domain Gaussian kernel function, and the pixel domain Gaussian kernel function, and calculating the filtered pixel values; and obtaining a low-illumination target image according to the filtered pixel values.

[0064] When configuring bilateral filtering parameters according to mine environment parameters, at least the neighborhood window is configured based on the mine environment parameters. The size of the spatial domain Gaussian kernel function , pixel domain Gaussian kernel function In this embodiment, the configuration of bilateral filtering parameters based on mine environment parameters is given as an example. For example, when the mine environment brightness is greater than >100 (cd / m²) and the mine dust concentration is <1000 (mg / m3), the neighborhood window The size of the spatial domain Gaussian kernel function can be 5 The width of the kernel function can be 3x3, the pixel domain Gaussian kernel function The width of the corresponding kernel function can be 3x3; in other cases, the neighborhood window The size of the spatial domain Gaussian kernel function can be 5 The width of the kernel function can be 9x9, the pixel domain Gaussian kernel function The width of the corresponding kernel function can be 9x9.

[0065] In this embodiment, for the spatial domain Gaussian kernel function, when the width of the kernel function is determined, the corresponding filtering processing using the spatial domain Gaussian kernel function can be determined.

[0066] Specifically, when the low-illumination source image is subjected to bilateral filtering processing, the calculation formula is: ; wherein, is the pixel value after filtering, is the pixel value in the original image, is the neighborhood window when the bilateral filtering processing is performed, is the normalized weight of the pixel value in the neighborhood window , the spatial domain Gaussian kernel function, is the pixel domain Gaussian kernel function.

[0067] Step 204, input the low-illumination target image into the DSC-ResNet network model, and extract the first image feature, the second image feature, the third image feature, and the fourth image feature through different sizes of convolution kernels.

[0068] The computer device can input the low-illumination target image into the image encoder module, and the encoder is a DSC-ResNet backbone network. The model structure of the DSC-ResNet network model is as shown in Figure 3 , which includes an embedding layer and a DSCBlock module, and is subjected to 7x7 convolution and then 3x3 maximum pooling operation.

[0069] ​Specifically, in one embodiment, the lightweight target detection method for small target detection in coal mine underground provided can further include a feature extraction process using a DSC-ResNet network model, and the specific process includes: inputting a low-illumination target image into the DSC-ResNet network model, performing convolution pooling operation through the encoder in the DSC-ResNet network model to generate preliminary features; according to the preliminary features, using different sizes of convolution kernels to perform spatial down-sampling and increase the number of channels on the low-illumination target image to obtain intermediate features; inputting the intermediate features into the DSC Block module to obtain image extraction features; obtaining first image features, second image features, third image features and fourth image features according to the image extraction features.

[0070] The computer device can input the low-illumination target image into the DSC-ResNet network model to gradually extract image features through four stages, and each stage contains an embedding layer and three DSC Block modules. At the beginning of the first stage, different sizes (1x1, 5x5, 7x7) of convolution kernels are used to perform spatial down-sampling and increase the number of channels on the image to generate three groups of intermediate features.

[0071] Then, the three groups of intermediate features can be input into the DSC Block module, so that the DSC Block module gradually extracts information from the image. Among them, a residual connection is arranged in each DSC Block module to reuse the input features, make up for the loss of inter-channel correlation information caused by the depth separation operation, down-sample and increase the number of channels. The number of feature map channels generated in each stage becomes twice the original, and the length and width are half of the previous stage. The DSC Block module can be described by the formula: ; wherein, is pixel-level addition.

[0072] ​​​In one embodiment, the lightweight target detection method for small target detection in a coal mine underground provided can further include a process of extracting image features of each stage respectively, and the specific process includes: inputting the image extraction features into a first size of convolution kernel after splicing, and performing pixel-level addition with the intermediate features to obtain first image features; performing convolution pooling operation on the first image features through an encoder, and performing spatial down-sampling to increase the number of channels, inputting the first image features into a second size of convolution kernel after processing through a DSC Block module to obtain second image features; performing convolution pooling operation on the second image features through an encoder, and performing spatial down-sampling to increase the number of channels, inputting the second image features into a third size of convolution kernel after processing through a DSC Block module to obtain third image features; performing convolution pooling operation on the third image features through an encoder, and performing spatial down-sampling to increase the number of channels, inputting the third image features into a fourth size of convolution kernel after processing through a DSC Block module to obtain fourth image features.

[0073] The computer device can splice the image extraction features in the channel dimension, then input 256 1x1 convolution kernels for convolution, and perform pixel-level addition with the generated intermediate features, and the first stage image features extracted are the first image features .

[0074] Then, the computer device can input the first image features into the second stage, extract image features through the DSC-ResNet network, splice the image extraction features in the channel dimension, then input 512 1x1 convolution kernels for convolution, and extract the second stage image features as the second image features .

[0075] The computer device can input the second image features into the third stage, extract image features through the DSC-ResNet network, splice the image extraction features in the channel dimension, then input 1024 1x1 convolution kernels for convolution, and extract the third stage image features as the third image features .

[0076] The computer device can input the third image features into the fourth stage, extract image features through the DSC-ResNet network, splice the image extraction features in the channel dimension, then input 2048 1x1 convolution kernels for convolution, and extract the fourth stage image features as the fourth image features .

[0077] In one embodiment, the DSC Block module includes a depth separable convolution layer, a standard convolution layer, a batch normalization layer, and an activation layer.

[0078] That is, each DSC Block module is composed of a DSConv (depth separable convolution) layer (convolution kernel size 3x3), two standard convolution layers (convolution kernel size 1x1), a batch normalization layer (BN), and an activation layer (ReLU).

[0079] Step 206, input the fourth image feature into the average pooling layer of the DSC-ResNet network model for pooling, and use the Softmax function for normalization processing to obtain linear one-dimensional data after processing.

[0080] Step 208, pixel-level addition is performed between the linear one-dimensional data and the position encoding vector to convert the first image sequence, and the first image sequence is input into the encoder for processing, and a feature map is generated through inverse linear projection.

[0081] The computer device can convert the feature map extracted by the DSC-ResNet network into a first image sequence through image embedding and position encoding operations, and input the converted first image sequence into the Transformer Encoder to obtain an encoded feature sequence, and generate a feature map F1 through inverse linear projection (through the transpose matrix The high-dimensional feature is mapped back to the original two-dimensional space.

[0082] Step 210, input the first image feature, the second image feature, the third image feature, the fourth image feature, and the feature map into the lightweight feature pyramid network model CA-LFPN for processing to obtain an output image.

[0083] Step 212, perform linearization operation on the output image, pixel-level addition with the linear one-dimensional data in the encoder, and input into the decoder, generate a prediction bounding box based on the output data of the decoder, and complete the target prediction.

[0084] Specifically, in one embodiment, the lightweight target detection method for small target detection in a coal mine provided can further include a process of processing features by a lightweight feature pyramid network model CA-LFPN, and the specific process includes: inputting the first image feature, the second image feature, the third image feature, the fourth image feature, and the feature map into the lightweight feature pyramid network model CA-LFPN, and obtaining five groups of features after processing by an attention mechanism module; performing convolution operation on the five groups of features, performing up-sampling processing on the fourth image feature, performing pixel-level addition between the up-sampled standard fourth image feature and the third image feature after the convolution operation to obtain a standard third image feature; performing up-sampling processing on the standard third image feature, and performing pixel-level addition between the up-sampled standard third image feature and the second image feature after the convolution operation to obtain a standard second image feature; performing up-sampling processing on the standard second image feature, and performing pixel-level addition between the up-sampled standard second image feature and the first image feature after the convolution operation to obtain a standard first image feature.

[0085] The computer device can input the generated feature map F1 and the four stage image features 、 、 、 into the lightweight feature pyramid network CA-LFPN at the same time. The CA-LFPN network model is as shown in Figure 4 , the feature map F1 and the four stage image features 、 、 、 are input, processed by an attention mechanism module CA, and then added after 1x1 and 3x3 convolution operations to obtain an output. Specifically, 、 、 、 and F1 are input into the CA attention mechanism module, and the obtained image features are added with the original input at the pixel level.

[0086] Then, the computer device can perform 1x1 convolution on the obtained five groups of features, i.e., the feature map F1 and the four stage image features 、 、 、 , perform 2x2 up-sampling on to make the image feature length and width of twice the original image, and then perform pixel-level addition between and after 1x1 convolution; then perform 2x2 up-sampling on the obtained 2x2 up-sampling is performed so that the image features are twice as long and twice as wide as the original image, and then 1x1 convolution is performed Pixel-level addition is performed; 3x3 convolution is performed on the five groups of features obtained, and then the channels are spliced.

[0087] In one embodiment, the lightweight target detection method for small target detection in a coal mine provided can further include a process of generating a predicted bounding box, and the specific process includes: after the standard first image features, the standard second image features, the standard third image features, and the standard fourth image features are subjected to convolution operation, a second image sequence is generated through linear projection; the first image sequence and the second image sequence are subjected to pixel-level addition, and the added image sequence is input into a decoder, and a predicted bounding box is generated based on the output data of the decoder.

[0088] The computer device can perform 3x3 convolution on the feature map processed by the lightweight feature pyramid network model CA-LFPN, and then generate a second image sequence through linear projection; the first image sequence and the second image sequence are subjected to pixel-level addition, and then input into a Transformer Dncoder decoder, and finally a predicted bounding box is generated.

[0089] The lightweight target detection method for small target detection in a coal mine provided in the present application first uses image enhancement technology to enhance the image, and then detects the small target image after enhancement, thereby improving the small target detection accuracy of the image obtained in the coal mine. Specifically, the DSC-ResNet network model is combined with the lightweight feature pyramid network model CA-LFPN, the low-level pixel-level features and the high-level semantic features are combined while realizing the lightweight of the model, the CA attention module is used to reduce the information loss to the minimum, the detection accuracy and the sensitivity to the new class position information are ensured under the condition of reducing the calculation cost, not only the high-level features with rich semantic information are extracted, but also the loss of low-level features providing accurate target position is avoided.

[0090] It should be understood that although each step in the above flowchart is displayed in sequence according to the direction of the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other orders. Moreover, at least part of the steps in the above flowchart can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily sequential, but can be executed in rotation or alternation with at least part of other steps or sub-steps or stages of other steps.

[0091] In one embodiment, as shown in Figure 5 A lightweight target detection system for small target detection in coal mine underground is provided, comprising: an image acquisition and processing module 510, a feature extraction module 520, a linear processing module 530, an encoding module 540, a feature processing module 550, and a prediction module 560, wherein:

[0092] The image acquisition and processing module 510 is configured to acquire a low-illumination source image of a small target in a coal mine underground, and acquire mine environment parameters. The bilateral filtering parameters are configured based on the mine environment parameters, and the low-illumination source image is processed based on the bilateral filtering parameters to obtain a low-illumination target image.

[0093] The feature extraction module 520 is configured to input the low-illumination target image into a DSC-ResNet network model, and extract a first image feature, a second image feature, a third image feature, and a fourth image feature through different sizes of convolution kernels.

[0094] The linear processing module 530 is configured to input the fourth image feature into an average pooling layer of the DSC-ResNet network model for pooling, and perform normalization processing using a Softmax function to obtain processed linear one-dimensional data.

[0095] The encoding module 540 is configured to add the linear one-dimensional data and the position encoding vector at the pixel level to convert the linear one-dimensional data into a first image sequence. The first image sequence is input into an encoder for processing, and a feature map is generated through inverse linear projection.

[0096] The feature processing module 550 is configured to input the first image feature, the second image feature, the third image feature, the fourth image feature, and the feature map into a lightweight feature pyramid network model CA-LFPN for processing to obtain an output image.

[0097] The prediction module 560 is configured to perform linearization operation on the output image, add the linear one-dimensional data in the encoder at the pixel level, input into a decoder, generate a prediction bounding box based on the output data of the decoder, and complete target prediction.

[0098] In one embodiment, the lightweight target detection system for small target detection in coal mine underground can be applied in a MFFL-DETR network model, as shown in Figure 6 The low-illumination target image after preprocessing is input into a DSC-ResNet network for feature extraction, pooled using an average pooling layer, normalized using a Softmax function, and then input into a lightweight feature pyramid network model CA-LFPN to complete target detection.

[0099] In one embodiment, the image acquisition and processing module 510 is further configured to acquire the ambient brightness of the coal mine underground through the illumination sensor, and measure the dust concentration of the coal mine underground through the dust sensor; take the ambient brightness and the dust concentration as the mine environment parameters, and configure the bilateral filtering parameters according to the mine environment parameters.

[0100] In one embodiment, the image acquisition and processing module 510 is further configured to determine a spatial neighborhood window of the low-illumination source image; acquire a normalized weight of pixel values in the spatial neighborhood window, determine a spatial domain Gaussian kernel function and a pixel domain Gaussian kernel function based on the spatial neighborhood window; perform bilateral filtering processing on the low-illumination source image according to the normalized weight, the spatial domain Gaussian kernel function and the pixel domain Gaussian kernel function, and calculate a filtered pixel value; and obtain a low-illumination target image according to the filtered pixel value.

[0101] In one embodiment, the feature extraction module 520 is further configured to input the low-illumination target image into the DSC-ResNet network model, perform convolutional pooling operation on the low-illumination target image through an encoder in the DSC-ResNet network model to generate preliminary features; use convolution kernels of different sizes to perform spatial down-sampling on the low-illumination target image and increase the number of channels according to the preliminary features to obtain intermediate features; input the intermediate features into a DSC Block module to obtain image extraction features; and obtain a first image feature, a second image feature, a third image feature and a fourth image feature according to the image extraction features.

[0102] In one embodiment, the feature extraction module 520 is further configured to input the image extraction features after being spliced into a convolution kernel of a first size, and add the image extraction features and the intermediate features at a pixel level to obtain the first image feature; perform convolutional pooling operation on the first image feature through an encoder, and perform spatial down-sampling to increase the number of channels, input the first image feature into a convolution kernel of a second size after being processed by the DSC Block module to obtain the second image feature; perform convolutional pooling operation on the second image feature through an encoder, and perform spatial down-sampling to increase the number of channels, input the second image feature into a convolution kernel of a third size after being processed by the DSC Block module to obtain the third image feature; perform convolutional pooling operation on the third image feature through an encoder, and perform spatial down-sampling to increase the number of channels, input the third image feature into a convolution kernel of a fourth size after being processed by the DSC Block module to obtain the fourth image feature.

[0103] In one embodiment, the DSC Block module includes a depth separable convolution layer, a standard convolution layer, a batch normalization layer and an activation layer.

[0104] In one embodiment, the prediction module 560 is further configured to input the first image feature, the second image feature, the third image feature, the fourth image feature and the feature map into a light-weight feature pyramid network model CA-LFPN, and obtain five groups of features after processing by an attention mechanism module; perform convolution operation on the five groups of features, perform up-sampling processing on the fourth image feature, perform pixel-level addition on the up-sampled standard fourth image feature and the third image feature after convolution operation, and obtain a standard third image feature; perform up-sampling processing on the standard third image feature, perform pixel-level addition on the up-sampled standard third image feature and the second image feature after convolution operation, and obtain a standard second image feature; perform up-sampling processing on the standard second image feature, perform pixel-level addition on the up-sampled standard second image feature and the first image feature after convolution operation, and obtain a standard first image feature.

[0105] In one embodiment, the prediction module 560 is further configured to generate a second image sequence by linear projection after performing convolution operation on the standard first image feature, the standard second image feature, the standard third image feature and the standard fourth image feature; perform pixel-level addition on the first image sequence and the second image sequence, and input the added image sequence into a decoder to generate a prediction bounding box based on output data of the decoder.

[0106] In one embodiment, a computer device, which can be a terminal, has an internal structure as shown in Figure 7 The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is configured to communicate with external terminals through network connection. The computer program is executed by the processor to implement a light-weight target detection method for small target detection in a coal mine underground. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad arranged on the shell of the computer device, or an external keyboard, touchpad or mouse, etc.

[0107] Those skilled in the art can understand that Figure 7 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0108] In one embodiment, a computer device is provided, comprising a memory and a processor, the memory storing a computer program, the processor implementing the steps of the lightweight target detection method for small target detection in coal mine underground when executing the computer program.

[0109] In one embodiment, a computer readable storage medium is provided, storing a computer program, the computer program implementing the steps of the lightweight target detection method for small target detection in coal mine underground when executed by a processor.

[0110] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware, and the computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, the processes of the above-mentioned embodiments can be included. Any reference to memory, storage, database or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM) and memory bus dynamic RAM (RDRAM) and the like.

[0111] Any combination of the technical features of the above embodiments can be made, and in order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combination of the technical features does not exist, it should be considered as the scope of the present application.

[0112] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are within the scope of the present application. Therefore, the scope of the patent of the present application should be subject to the appended claims.

Claims

1. A lightweight target detection method for small target detection in coal mines, characterized in that: The method comprises: Acquire a collected low-illumination source image of a small target in a coal mine, collect mine environmental parameters, configure bilateral filtering parameters based on the mine environmental parameters, and process the low-illumination source image based on the bilateral filtering parameters to obtain a low-illumination target image; Inputting the low-light target image into the DSC-ResNet network model, and extracting the first image feature, the second image feature, the third image feature, and the fourth image feature respectively through convolution kernels of different sizes; Inputting the fourth image feature into the average pooling layer of the DSC-ResNet network model for pooling, and performing normalization processing using a Softmax function to obtain processed linear one-dimensional data; Performing pixel-level addition on the linear one-dimensional data and the position encoding vector to convert the data into a first image sequence, inputting the first image sequence into an encoder for processing, and generating a feature map through inverse linear projection; Inputting the first image feature, the second image feature, the third image feature, the fourth image feature, and the feature map into a lightweight feature pyramid network model CA-LFPN for processing to obtain an output image; The output image is linearized, pixel-wise added to the linear one-dimensional data in the encoder, and input into the decoder. A prediction bounding box is generated based on the output data of the decoder to complete target prediction.

2. The lightweight target detection method for small target detection in coal mines according to claim 1, characterized in that: Collecting mine environmental parameters and configuring bilateral filtering parameters based on the mine environmental parameters, including: The ambient brightness of the coal mine is collected through the light sensor, and the dust concentration in the coal mine is measured through the dust sensor; The ambient brightness and dust concentration are used as mine environmental parameters, and bilateral filtering parameters are configured according to the mine environmental parameters.

3. The lightweight target detection method for small target detection in coal mines according to claim 1, characterized in that: The low-illumination source image is processed based on the bilateral filtering parameters to obtain a low-illumination target image, comprising: Determining a spatial neighborhood window of the low-illumination source image; Obtaining normalized weights of pixel values ​​within the spatial neighborhood window, and determining a spatial domain Gaussian kernel function and a pixel domain Gaussian kernel function based on the spatial neighborhood window; Performing bilateral filtering on the low-illumination source image according to the normalization weight, the spatial domain Gaussian kernel function, and the pixel domain Gaussian kernel function to calculate filtered pixel values; A low-illumination target image is obtained according to the filtered pixel values.

4. The lightweight target detection method for small target detection in coal mines according to claim 1, characterized in that: The low-light target image is input into the DSC-ResNet network model, and the first image feature, the second image feature, the third image feature, and the fourth image feature are extracted respectively through convolution kernels of different sizes, including: Inputting the low-light target image into the DSC-ResNet network model, performing convolution and pooling operations through the encoder in the DSC-ResNet network model to generate preliminary features; According to the preliminary features, using convolution kernels of different sizes to spatially downsample the low-light target image and increase the number of channels to obtain intermediate features; Input the intermediate features into the DSC Block module to obtain image extraction features; A first image feature, a second image feature, a third image feature, and a fourth image feature are obtained according to the image extraction features.

5. The lightweight target detection method for small target detection in coal mines according to claim 4, characterized in that: Extracting features from the image to obtain a first image feature, a second image feature, a third image feature, and a fourth image feature includes: The image extracted features are spliced ​​and input into a convolution kernel of a first size, and pixel-wise added to the intermediate features to obtain a first image feature; The first image feature is convolutionally pooled through the encoder, spatially downsampled to increase the number of channels, and then processed by the DSC Block module and input into the convolution kernel of the second size to obtain the second image feature; The second image feature is convolutionally pooled through the encoder, spatially downsampled to increase the number of channels, and then processed by the DSC Block module and input into the convolution kernel of the third size to obtain the third image feature; The third image feature is convolutionally pooled through the encoder and spatially downsampled to increase the number of channels. After being processed by the DSC Block module, it is input into the convolution kernel of the fourth size to obtain the fourth image feature.

6. The lightweight target detection method for small target detection in coal mines according to claim 5, characterized in that: The DSC Block module includes a depth-wise separable convolution layer, a standard convolution layer, a batch normalization layer, and an activation layer.

7. The lightweight target detection method for small target detection in coal mines according to claim 1, characterized in that: Inputting the first image feature, the second image feature, the third image feature, the fourth image feature, and the feature map into a lightweight feature pyramid network model CA-LFPN for processing, including: The first image feature, the second image feature, the third image feature, the fourth image feature, and the feature map are input into the lightweight feature pyramid network model CA-LFPN, and five sets of features are obtained after processing by the attention mechanism module; Performing a convolution operation on the five sets of features, and performing upsampling on the fourth image feature, and performing pixel-level addition on the upsampled standard fourth image feature and the convolution-operated third image feature to obtain a standard third image feature; Performing upsampling processing on the standard third image feature and performing pixel-level addition on the second image feature after the convolution operation to obtain the standard second image feature; The standard second image feature is up-sampled and added to the first image feature after the convolution operation at a pixel level to obtain the standard first image feature.

8. The lightweight target detection method for small target detection in coal mines according to claim 7, characterized in that: Performing a linearization operation on the output image, performing pixel-level addition on the linear one-dimensional data in the encoder, inputting the output image into a decoder, and generating a predicted bounding box based on the output data of the decoder, including: After performing a convolution operation on the standard first image feature, the standard second image feature, the standard third image feature, and the standard fourth image feature, a second image sequence is generated by linear projection; The first image sequence and the second image sequence are added at a pixel level, and the added image sequence is input into a decoder, and a prediction bounding box is generated based on output data of the decoder.

9. A lightweight target detection system for detecting small targets in coal mines, characterized in that: The system comprises: An image acquisition and processing module is used to acquire a low-illumination source image of a small target in a coal mine, collect mine environmental parameters, configure bilateral filtering parameters based on the mine environmental parameters, and process the low-illumination source image based on the bilateral filtering parameters to obtain a low-illumination target image; A feature extraction module is used to input the low-light target image into the DSC-ResNet network model and extract the first image feature, the second image feature, the third image feature, and the fourth image feature respectively through convolution kernels of different sizes; a linear processing module, configured to input the fourth image feature into the average pooling layer of the DSC-ResNet network model for pooling, and perform normalization processing using a Softmax function to obtain processed linear one-dimensional data; an encoding module, configured to perform pixel-level addition of the linear one-dimensional data and the position encoding vector to convert the data into a first image sequence, input the first image sequence into an encoder for processing, and generate a feature map through inverse linear projection; a feature processing module, configured to input the first image feature, the second image feature, the third image feature, the fourth image feature, and the feature map into a lightweight feature pyramid network model CA-LFPN for processing to obtain an output image; A prediction module is used to perform a linearization operation on the output image, perform pixel-level addition on the linear one-dimensional data in the encoder, input the result to the decoder, generate a prediction bounding box based on the output data of the decoder, and complete the target prediction.