Under-mine target detection method based on multi-scale feature fusion

Through a multi-scale feature fusion method based on deformable convolution and cross-attention mechanism, the problem of insufficient feature extraction and fusion in underground target detection is solved, and high-precision underground target detection is achieved, especially the effective identification of miners and their clothing.

CN120635636APending Publication Date: 2025-09-12DALIAN MARITIME UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510583982.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing underground target detection methods lack the ability to adaptively model spatial geometric structures in the feature extraction stage. Feature fusion strategies ignore detailed information, and traditional regression loss functions cannot effectively distinguish the quality of detection frames, resulting in low detection accuracy and efficiency.

Method used

A feature extraction network based on deformable convolution and cross attention mechanism is adopted, combined with multi-scale feature aggregation and IoU loss function, to perform detection box regression and candidate box screening, and the object detection model is trained using the coal mine image dataset.

Benefits of technology

It improves the accuracy and efficiency of underground target detection, especially the detection of miners and their clothing far away from surveillance cameras, and can generate high-precision detection frame categories and location information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635636A_ABST
    Figure CN120635636A_ABST
Patent Text Reader

Abstract

The invention provides an under-mine target detection method based on multi-scale feature fusion, and the method comprises the steps: obtaining an under-mine image, carrying out the shape preprocessing of the under-mine image, obtaining an image with a fixed size, and taking the image as an input image I; constructing a feature extraction network based on deformable convolution and a cross attention mechanism, and performing feature extraction on the input image I to obtain four output feature maps P2, P3, P4 and P5 of different scales; constructing a feature aggregation network based on attention and scale consistency, and performing feature aggregation on the feature maps P2, P3, P4 and P5 to obtain three feature maps F3, F4 and F5; constructing an IoU loss function based on spatial perception, and performing detection frame regression on the feature maps F3, F4 and F5 to obtain a plurality of candidate detection frame sets B = {B1,..., Bi,... BN}; and based on linear attenuation, performing target positioning in a candidate box screening stage to obtain category information and position coordinates of the miner and the wear thereof, and performing equal-proportion coordinate change on the original image according to the position coordinates to obtain actual coordinate values.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of underground target detection, and in particular to an underground target detection method based on multi-scale feature fusion. Background Art

[0002] Existing methods for detecting targets in underground mines primarily use a single-stage target detection algorithm. This method bypasses the secondary classification and selection of candidate anchor boxes and instead performs regression directly on the candidate anchor boxes, effectively improving detection speed. This method extracts image features through multiple convolutional feature extraction networks, performs feature fusion, and ultimately uses the detection head to regress the target category and detection box coordinates.

[0003] There are some problems with the existing underground target detection, as follows:

[0004] (1) Existing underground target detection networks generally use standard convolution operations in the feature extraction stage. Their receptive fields and convolution kernel sampling positions are relatively fixed, and they lack the ability to adaptively model spatial geometric structures. This fixed sampling mechanism makes it difficult to achieve effective shape alignment and feature matching when facing targets with irregular shapes and variable postures in underground mine environments. This results in insufficient expressiveness of the extracted features, which in turn affects the accuracy of target positioning and recognition.

[0005] (2) Existing underground target detection networks have limitations in the design of feature fusion strategies. They usually only focus on simple splicing or weighted fusion between features, ignoring the information loss that may occur during the transmission and calculation of detailed features. In particular, they do not adequately preserve key texture and edge information in low-level features. At the same time, existing methods lack effective expression and modeling mechanisms for high-resolution, large-scale feature maps, resulting in limited performance when detecting small targets or targets with complex structures.

[0006] (3) The traditional regression-based detection box loss function lacks effective distinction of detection box quality during training and cannot accurately determine which detection boxes are high-quality candidates. Therefore, in the early stages of model training, the network will indiscriminately perform regression operations on a large number of low-quality detection boxes, which not only reduces training efficiency but also weakens the model's focus on high-quality detection boxes, thereby affecting the final positioning accuracy and convergence speed.

[0007] Therefore, in the underground target detection algorithm, in-depth research and optimization are needed for the above problems to improve the accuracy and practicality of the algorithm. Summary of the Invention

[0008] To address the aforementioned technical issues, a method for detecting underground targets based on multi-scale feature fusion is provided. This method utilizes a coal mine image dataset to train a target detection model. This model effectively improves detection of miners and their clothing located far from surveillance cameras, generating corresponding detection frame categories and location information.

[0009] The technical means adopted in the present invention are as follows:

[0010] A method for detecting underground targets based on multi-scale feature fusion, comprising:

[0011] S1. Acquire an image of a mine shaft, perform shape preprocessing on the image, obtain an image of a fixed size, and use the image as an input image I;

[0012] S2. Based on deformable convolution and cross attention mechanism, a feature extraction network is constructed to extract features from the input image I and obtain four output feature maps P2, P3, P4 and P5 at different scales;

[0013] S3. Based on attention and scale consistency, a feature aggregation network is constructed to aggregate the feature maps P2, P3, P4, and P5 to obtain three feature maps F3, F4, and F5.

[0014] S4, based on spatial perception, construct the IoU loss function, perform detection frame regression on the feature maps F3, F4 and F5, and obtain multiple candidate detection frame sets B = {B1,...,B i ,...B N};

[0015] S5. Based on linear attenuation, target positioning is performed in the candidate frame screening stage to obtain the category information and position coordinates of the miners and their wearing items. The actual coordinate values ​​are obtained by performing proportional coordinate changes on the original image based on the position coordinates.

[0016] Furthermore, step S2 specifically includes:

[0017] S21, after the input image I passes through the convolution layer, the height H and width W are both changed to 160. After passing through the channel transformation layer (C2f), the height H and width W of the feature map are not changed, and the feature map is transformed in the channel dimension to output the feature map.

[0018] After S22 and feature map P2 pass through the convolution layer and channel transformation layer (C2f), the height H and width W of the feature map are changed to 80, and the output feature map

[0019] After S23 and feature map P3 pass through the convolution layer and channel feature mapping module, the height H and width W are both changed to 40, and the output feature map

[0020] After S24 and feature map P4 pass through the convolution layer, the height H and width W are both changed to 20, and then pass through the channel feature mapping module and fast spatial pyramid pooling to output the feature map

[0021] Furthermore, the channel feature mapping module divides the input feature map into feature maps through Split, passes through n bottleneck feature extraction modules, uses Concat to connect the previous intermediate feature maps in the output stage, and then passes through the convolution layer to obtain the final output feature map.

[0022] Furthermore, the bottleneck feature extraction module performs preliminary feature extraction on the input features through the convolution layer, and obtains the feature map after learning the offset through deformable convolution (DCNv2), wherein the calculation formula of deformable convolution (DCNv2) is as follows:

[0023]

[0024] in, is the input feature map, is the output feature map, K is the number of convolution kernel sampling points, W k Represents the learnable weight of the k-th sampling point of the convolution kernel, P k Represents the relative offset of the kth sampling point of the convolution kernel, P represents the position coordinate of the input feature map, ΔP k represents the learnable offset of the kth sampling position, X(P) represents the input feature Figure X The eigenvector at position P, X(P+P k +ΔP k ) represents the input feature Figure X The eigenvector after the position is shifted, Δm k ∈[0,1] is a learnable modulation factor, which is used to adjust the weight of the k-th sampling point; the output feature at the P position is finally obtained

[0025] Furthermore, the feature map after learning the offset is subjected to the cross-attention mechanism to obtain an output feature map. The feature map output by the cross-attention mechanism is residually connected with the original input feature map to obtain the final output feature map, as follows:

[0026] The input feature map divides the features into subgroups along the channel dimension g represents the g-th channel subgroup;

[0027] Subgroup X g The left feature is obtained after 1×1 convolution branch and normalization Subgroup Xg Then pass through the 3×3 convolution branch to get the right feature

[0028] The left feature L g With the right feature R g By flattening the spatial dimension into L′ g and R′ g , The cross attention calculation is completed through the matrix multiplication step, and the calculation formula is as follows:

[0029] W g =σ(α g L′ g +β g R′ g )

[0030]

[0031] in, is the attention feature vector of the left 1×1 convolution branch, is the attention feature vector of the right 3×3 convolution branch, σ is the Sigmoid activation function, is the weighted subgroup feature, is the overall output feature after splicing.

[0032] Furthermore, step S3 specifically includes:

[0033] S31. For each feature map F i , the multi-scale feature aggregation module performs channel alignment operation, and the calculation process is as follows:

[0034]

[0035] Among them, W is the weight matrix, b is the convolution bias matrix, is the channel scaling factor, b BN is the normalized bias, δ represents the ReLU activation function, Represents the feature map after channel alignment;

[0036] S32, for the feature map F with a shape size of 0.5C×2H×2W L Perform channel compression to achieve information compression and obtain the compressed feature map

[0037] S33, compressed feature map Performing maximum pooling and average pooling operations separately to extract significant local information while retaining global background and boundary features, and then fusing the two types of information by element-by-element addition, thus taking into account both details and overall perception;

[0038] S34, for the feature map F with a shape size of C×H×W M Perform channel alignment to obtain feature maps

[0039] S35, for the feature map F with a shape size of 2C×0.5H×0.5W S Upsampling and feature map after channel alignment The shape remains consistent;

[0040] S36. Concat the three feature maps of different scales after spatial size alignment to obtain a fused feature map with both detail features and spatial position information. After passing through the feature aggregation network, three feature maps F3, F4 and F5 are obtained.

[0041] Furthermore, step S4 specifically includes:

[0042] S41, transform the feature map into and

[0043] S42. Add 5 channels in the channel dimension, of which 4 channels are used to predict the position coordinates of the detection box, and 1 channel is used to determine whether there is an object at that position;

[0044] S43, each pixel point spatial position passes through the multi-layer perceptron, and the number of channels is mapped to the probability of each category. The first 4 are [t x ,t y ,t w ,t h ], the mapping coordinates are calculated by the decoding formula, the decoding formula is as follows:

[0045] x=s·(2σ(t x )-0.5+i)

[0046] y=s·(2σ(t y )-0.5+j)

[0047] w=s·(2σ(t w )) 2

[0048] h=s·(2σ(t h )) 2

[0049] Among them, s is the downsampling multiple, σ is the Sigmoid activation function, (i, j) is the feature map position of the prediction point, t x With t y Represents the position prediction value, (x, y) is the position of the prediction box in the input image, t w With th are the predicted values ​​of width and height, respectively, w and h are the width and height of the predicted box in the input image;

[0050] S44. The detection head performs detection box regression through coordinate mapping. This process uses a loss function to constrain it. The IoU loss function based on spatial perception is as follows:

[0051]

[0052] l WIoU =r WIoU (1-IoU)

[0053]

[0054] Among them, α is a positive balance parameter, v is used to measure the consistency of aspect ratio, and w gt With h gt is the width and height of the real detection frame, w and h are the width and height of the output detection frame, c is the diagonal length of the minimum circumscribed rectangle, b is the output detection frame position, b gt is the true detection box position, ρ(b,b gt ) represents the Euclidean distance between the center points, W g and H g is the width and height of the minimum bounding rectangle of the real detection box and the output detection box, in r WIoU In order to prevent the gradient from converging during the back propagation of the function, Separated from the computational graph and does not participate in back propagation;

[0055] S45, after regression, multiple candidate detection frame sets B={B1,...,B i ,...B N}, where N represents the number of categories.

[0056] Furthermore, step S5 specifically includes:

[0057] S51, according to the confidence threshold t, the i-th type detection box set B i Perform preliminary screening to obtain the detection frame group G i , G i ={g i1 ,...,g ij ,...,g iM}, where M represents the number of detection frames after preliminary screening, and initializes the output sequence D i , at this time D i is an empty set;

[0058] S52, traverse the detection frame group G i, find the detection box M with the highest confidence, and add the detection box M to the output sequence D i ;

[0059] S53, based on the detection frame M and the detection frame group G i The rest of the detection boxes g ij , calculate the intersection-over-union ratio, the calculation formula is as follows:

[0060]

[0061] Among them, IoU ij Represents the detection box M and the detection box group G i The j-th detection box g in ij The intersection and union ratio of

[0062] S54, remove the detection frame M from the detection frame group G i Eliminate it and only calculate each intersection over IoU ij The confidence of boxes greater than or equal to the threshold is linearly decayed, and the calculation formula is as follows:

[0063]

[0064] Among them, τ is the intersection-over-union ratio threshold, s ij is the confidence score of the j-th detection box of the i-th category, s′ ij is the confidence score after transformation;

[0065] S55, from the detection frame group G i Remove the low confidence score s′ ij The corresponding detection frame, until the detection frame group G i There are no remaining candidate boxes, and the output D is obtained i ;

[0066] S56: The remaining categories also perform the process from step S51 to step S55 to obtain the final output D={D1,...,D i ,...,D N}.

[0067] Compared with the prior art, the present invention has the following advantages:

[0068] 1. The present invention provides a method for underground target detection based on multi-scale feature fusion, which can use coal mine image datasets to train a target detection model. This model can effectively improve the detection effect of miners and their clothing far away from surveillance cameras, and generate corresponding detection frame categories and location information.

[0069] 2. The present invention provides a method for detecting underground targets based on multi-scale feature fusion, which only needs to be trained using the existing underground target annotated dataset to obtain a high-precision underground target detection inference model.

[0070] Based on the above reasons, the present invention can be widely promoted in fields such as underground target detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0072] Figure 1 Flow chart of the method of the present invention.

[0073] Figure 2 This is the feature extraction network structure diagram of the present invention.

[0074] Figure 3 This is a structural diagram of the channel feature mapping module of the present invention.

[0075] Figure 4 This is the structural diagram of the bottleneck feature extraction module of the present invention.

[0076] Figure 5 This is the structural diagram of the cross-attention mechanism of the present invention.

[0077] Figure 6 This is a diagram of the feature aggregation network structure of the present invention.

[0078] Figure 7 This is a structural diagram of the multi-scale feature splicing module of the present invention.

[0079] Figure 8 This is a flow chart of the positioning of the i-th type of target based on linear attenuation in the present invention.

[0080] Figure 9 This is a target detection result diagram provided by an embodiment of the present invention.

[0081] Figure 10 This is a curve diagram of the detection box loss change provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0082] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0083] It should be noted that the terms "including" and "having" and any variations thereof in the specification and claims of the present invention and the above-mentioned drawings are intended to cover non-exclusive inclusions. For example, a process, method, system, product or apparatus comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or apparatuses.

[0084] like Figure 1 As shown, the present invention provides a method for detecting underground targets based on multi-scale feature fusion, comprising:

[0085] S1. Acquire an image of a mine shaft, perform shape preprocessing on the image, obtain an image of a fixed size, and use the image as an input image I;

[0086] S2. Based on deformable convolution and cross attention mechanism, a feature extraction network is constructed to extract features from the input image I and obtain four output feature maps P2, P3, P4 and P5 at different scales;

[0087] S3. Based on attention and scale consistency, a feature aggregation network is constructed to aggregate the feature maps P2, P3, P4, and P5 to obtain three feature maps F3, F4, and F5.

[0088] S4, based on spatial perception, construct the IoU loss function, perform detection frame regression on the feature maps F3, F4 and F5, and obtain multiple candidate detection frame sets B = {B1,...,B i ,...B N};

[0089] S5. Based on linear attenuation, target positioning is performed in the candidate frame screening stage to obtain the category information and position coordinates of the miners and their wearing items. The actual coordinate values ​​are obtained by performing proportional coordinate changes on the original image based on the position coordinates.

[0090] In specific implementation, as a preferred embodiment of the present invention, step S2 specifically includes:

[0091] S21, feature extraction network structure based on deformable convolution and cross attention mechanism Figure 2As shown, after the input image I passes through the convolution layer, the height H and width W both become 160. After passing through the channel transformation layer (C2f), the height H and width W of the feature map are not changed, and the feature map is transformed in the channel dimension to output the feature map.

[0092] After S22 and feature map P2 pass through the convolution layer and channel transformation layer (C2f), the height H and width W of the feature map are changed to 80, and the output feature map

[0093] After S23 and feature map P3 pass through the convolution layer and channel feature mapping module, the height H and width W are both changed to 40, and the output feature map

[0094] After S24 and feature map P4 pass through the convolution layer, the height H and width W are both changed to 20, and then pass through the channel feature mapping module and fast spatial pyramid pooling to output the feature map

[0095] In specific implementation, as a preferred embodiment of the present invention, the structure of the channel feature mapping module is as follows: Figure 3 As shown in the figure, after the input feature map is divided into feature maps by Split, it passes through n bottleneck feature extraction modules, and uses Concat to connect the previous intermediate feature maps in the output stage, and then passes through the convolution layer to obtain the final output feature map.

[0096] In specific implementation, as a preferred embodiment of the present invention, the structure of the bottleneck feature extraction module is as follows: Figure 4 As shown in the figure, the input features are subjected to preliminary feature extraction through the convolution layer, and the feature map after learning the offset is obtained through deformable convolution (DCNv2). The calculation formula of deformable convolution (DCNv2) is as follows:

[0097]

[0098] in, is the input feature map, is the output feature map, K is the number of convolution kernel sampling points, W k Represents the learnable weight of the k-th sampling point of the convolution kernel, P k Represents the relative offset of the kth sampling point of the convolution kernel, P represents the position coordinate of the input feature map, ΔP k represents the learnable offset of the kth sampling position, X(P) represents the input feature Figure X The eigenvector at position P, X(P+P k +ΔP k ) represents the input feature Figure X The eigenvector after the position is shifted, Δmk ∈[0,1] is a learnable modulation factor, which is used to adjust the weight of the k-th sampling point; the output feature at the P position is finally obtained

[0099] In specific implementation, as a preferred embodiment of the present invention, the feature map after learning the offset is subjected to a cross-attention mechanism to obtain an output feature map, and the feature map output by the cross-attention mechanism is residually connected with the original input feature map to obtain the final output feature map, as follows:

[0100] The cross attention mechanism structure is as follows Figure 5 As shown, the input feature map divides the features into subgroups in the channel dimension g represents the g-th channel subgroup; subgroup X g The left feature is obtained after 1×1 convolution branch and normalization Subgroup X g Then pass through the 3×3 convolution branch to get the right feature The left feature L g With the right feature R g By flattening the spatial dimension into L′ g and The cross attention calculation is completed through the matrix multiplication step, and the calculation formula is as follows:

[0101] W g =σ(α g L′ g +β g R′ g )

[0102]

[0103] in, is the attention feature vector of the left 1×1 convolution branch, is the attention feature vector of the right 3×3 convolution branch, σ is the Sigmoid activation function, is the weighted subgroup feature, is the overall output feature after splicing.

[0104] In specific implementation, as a preferred embodiment of the present invention, step S3 specifically includes:

[0105] S31. For each feature map F i , the multi-scale feature aggregation module performs channel alignment operation, and the calculation process is as follows:

[0106]

[0107] Among them, W is the weight matrix, b is the convolution bias matrix, is the channel scaling factor, b BN is the normalized bias, δ represents the ReLU activation function, Represents the feature map after channel alignment;

[0108] S32, for the feature map F with a shape size of 0.5C×2H×2W L Perform channel compression to achieve information compression and obtain the compressed feature map

[0109] S33, compressed feature map Performing maximum pooling and average pooling operations separately to extract significant local information while retaining global background and boundary features, and then fusing the two types of information by element-by-element addition, thus taking into account both details and overall perception;

[0110] S34, for the feature map F with a shape size of C×H×W M Perform channel alignment to obtain feature maps

[0111] S35, for the feature map F with a shape size of 2C×0.5H×0.5W S Upsampling and feature map after channel alignment The shape remains consistent;

[0112] S36. Concat the three feature maps of different scales after spatial size alignment to obtain a fused feature map with both detail features and spatial position information. After passing through the feature aggregation network, three feature maps F3, F4 and F5 are obtained.

[0113] In this embodiment, the feature aggregation network structure based on attention and scale consistency is as follows Figure 6 As shown in Figure 2, the feature map obtained in step S2 is aggregated through three multi-scale feature splicing modules. The multi-scale feature splicing module structure in the feature aggregation network structure is shown in Figure 2. Figure 7 As shown in Figure 3, this module aggregates feature maps of different resolutions and uses cross-layer feature connections to effectively integrate multi-scale information.

[0114] In specific implementation, as a preferred embodiment of the present invention, step S4 specifically includes:

[0115] S41, transform the feature map into and

[0116] S42. Add 5 channels in the channel dimension, of which 4 channels are used to predict the position coordinates of the detection box, and 1 channel is used to determine whether there is an object at that position;

[0117] S43, each pixel point spatial position passes through the multi-layer perceptron, and the number of channels is mapped to the probability of each category. The first 4 are [t x ,t y ,t w ,t h ], the mapping coordinates are calculated by the decoding formula, the decoding formula is as follows:

[0118] x=s·(2σ(t x )-0.5+i)

[0119] y=s·(2σ(t y )-0.5+j)

[0120] w=s·(2σ(t w )) 2

[0121] h=s·(2σ(t h )) 2

[0122] Among them, s is the downsampling multiple, σ is the Sigmoid activation function, (i, j) is the feature map position of the prediction point, t x With t y Represents the position prediction value, (x, y) is the position of the prediction box in the input image, t w With t h are the predicted values ​​of width and height, respectively, w and h are the width and height of the predicted box in the input image;

[0123] S44. The detection head performs detection box regression through coordinate mapping. This process uses a loss function to constrain it. The IoU loss function based on spatial perception is as follows:

[0124]

[0125] l WIoU =r WIoU (1-IoU)

[0126]

[0127] Among them, α is a positive balance parameter, v is used to measure the consistency of aspect ratio, and w gt With h gt is the width and height of the real detection frame, w and h are the width and height of the output detection frame, c is the diagonal length of the minimum circumscribed rectangle, b is the output detection frame position, b gt is the true detection box position, ρ(b,b gt ) represents the Euclidean distance between the center points, W g and H gis the width and height of the minimum bounding rectangle of the real detection box and the output detection box, in r WIoU In order to prevent the gradient from converging during the back propagation of the function, Separated from the computational graph and does not participate in back propagation;

[0128] S45, after regression, multiple candidate detection frame sets B={B1,...,B i ,...B N}, where N represents the number of categories.

[0129] In specific implementation, as a preferred embodiment of the present invention, in step S5, target positioning is performed in the candidate frame screening stage, and the target positioning processing flow of the i-th category is as follows: Figure 8 As shown, specifically including:

[0130] S51, according to the confidence threshold t, the i-th type detection box set B i Perform preliminary screening to obtain the detection frame group G i , G i ={g i1 ,...,g ij ,...,g iM}, where M represents the number of detection frames after preliminary screening, and initializes the output sequence D i , at this time D i is an empty set;

[0131] S52, traverse the detection frame group G i , find the detection box M with the highest confidence, and add the detection box M to the output sequence D i ;

[0132] S53, based on the detection frame M and the detection frame group G i The rest of the detection boxes g ij , calculate the intersection-over-union ratio, the calculation formula is as follows:

[0133]

[0134] Among them, IoU ij Represents the detection box M and the detection box group G i The j-th detection box g in ij The intersection and union ratio of

[0135] S54, remove the detection frame M from the detection frame group G i Eliminate it and only calculate each intersection over IoU ij The confidence of boxes greater than or equal to the threshold is linearly decayed, and the calculation formula is as follows:

[0136]

[0137] Among them, τ is the intersection-over-union ratio threshold, s ij is the confidence score of the j-th detection box of the i-th category, s′ ij is the confidence score after transformation;

[0138] S55, from the detection frame group G i Remove the low confidence score s′ ij The corresponding detection frame, until the detection frame group G i There are no remaining candidate boxes, and the output D is obtained i ;

[0139] S56: The remaining categories also perform the process from step S51 to step S55 to obtain the final output D={D1,...,D i ,...,D N}.

[0140] Comparative Example

[0141] The mine target detection method proposed by the present invention can effectively detect small-sized targets under the mine. Figure 9 As shown; the loss function change curve during the detection box regression process is shown in Figure 10. The blue method proposed in this invention has a lower detection box loss under the same epoch.

[0142] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for detecting underground targets based on multi-scale feature fusion, characterized in that: include: S1. Acquire an image of a mine shaft, perform shape preprocessing on the image, obtain an image of a fixed size, and use the image as an input image I; S2. Based on deformable convolution and cross attention mechanism, a feature extraction network is constructed to extract features from the input image I and obtain four output feature maps P2, P3, P4 and P5 at different scales; S3. Based on attention and scale consistency, a feature aggregation network is constructed to aggregate the feature maps P2, P3, P4, and P5 to obtain three feature maps F3, F4, and F5. S4, based on spatial perception, construct the IoU loss function, perform detection frame regression on the feature maps F3, F4 and F5, and obtain multiple candidate detection frame sets B = {B1,...,B i ,...B N }; S5. Based on linear attenuation, target positioning is performed in the candidate frame screening stage to obtain the category information and position coordinates of the miners and their wearing items. The actual coordinate values ​​are obtained by performing proportional coordinate changes on the original image based on the position coordinates.

2. The method for detecting underground targets based on multi-scale feature fusion according to claim 1, characterized in that: Step S2 specifically includes: S21, after the input image I passes through the convolution layer, the height H and width W are both changed to 160. After passing through the channel transformation layer, the height H and width W of the feature map are not changed, and the feature map is transformed in the channel dimension to output the feature map P2. After S22 and feature map P2 pass through the convolution layer and channel transformation layer, the height H and width W of the feature map are changed to 80, and the output feature map P3 is After S23 and the feature map P3 pass through the convolution layer and the channel feature mapping module, the height H and width W are both changed to 40, and the output feature map P4 is After S24 and feature map P4 pass through the convolution layer, the height H and width W are both changed to 20, and then pass through the channel feature mapping module and fast spatial pyramid pooling to output feature map P5.

3. The method for detecting underground targets based on multi-scale feature fusion according to claim 2, characterized in that: The channel feature mapping module divides the input feature map into feature maps through Split, passes through n bottleneck feature extraction modules, uses Concat to connect the previous intermediate feature maps in the output stage, and then passes through the convolution layer to obtain the final output feature map.

4. The method for detecting underground targets based on multi-scale feature fusion according to claim 3, characterized in that: The bottleneck feature extraction module performs preliminary feature extraction on the input features through the convolution layer and obtains the feature map after learning the offset through deformable convolution. The calculation formula of deformable convolution is as follows: in, is the input feature map, is the output feature map, K is the number of convolution kernel sampling points, W k Represents the learnable weight of the k-th sampling point of the convolution kernel, P k Represents the relative offset of the kth sampling point of the convolution kernel, P represents the position coordinate of the input feature map, ΔP k represents the learnable offset of the kth sampling position, X(P) represents the feature vector at position P in the input feature map X, and X(P+P k +ΔP k ) represents the feature vector after the input feature map X is offset to the position, Δm k ∈[0,1] is a learnable modulation factor, which is used to adjust the weight of the k-th sampling point; the output feature at the P position is finally obtained 5. The method for detecting underground targets based on multi-scale feature fusion according to claim 4, characterized in that: The feature map after learning the offset is subjected to the cross attention mechanism to obtain the output feature map. The feature map after the cross attention mechanism output is residually connected with the original input feature map to obtain the final output feature map, which is as follows: The input feature map divides the features into subgroups X along the channel dimension g , g represents the g-th channel subgroup; Subgroup X g After 1×1 convolution branch and normalization, the left feature L is obtained g , Subgroup X g Then, through the 3×3 convolution branch, we get the right feature R g , The left feature L g With the right feature R g By flattening the spatial dimension into L′ g and R′ g , The cross attention calculation is completed through the matrix multiplication step, and the calculation formula is as follows: W g =σ(α g L′ g +b g R′ g ) in, is the attention feature vector of the left 1×1 convolution branch, is the attention feature vector of the right 3×3 convolution branch, σ is the Sigmoid activation function, is the weighted subgroup feature, is the overall output feature after splicing.

6. The method for detecting underground targets based on multi-scale feature fusion according to claim 1, characterized in that: Step S3 specifically includes: S31. For each feature map F i , the multi-scale feature aggregation module performs channel alignment operation, and the calculation process is as follows: Among them, W is the weight matrix, b is the convolution bias matrix, is the channel scaling factor, b BN is the normalized bias, δ represents the ReLU activation function, Represents the feature map after channel alignment; S32, for the feature map F with a shape size of 0.5C×2H×2W L Perform channel compression to achieve information compression and obtain the compressed feature map S33, compressed feature map Performing maximum pooling and average pooling operations separately to extract significant local information while retaining global background and boundary features, and then fusing the two types of information by element-by-element addition, thus taking into account both details and overall perception; S34, for the feature map F with a shape size of C×H×W M Perform channel alignment to obtain feature maps S35, for the feature map F with a shape size of 2C×0.5H×0.5W S Upsampling and feature map after channel alignment The shape remains consistent; S36. Concat the three feature maps of different scales after spatial size alignment to obtain a fused feature map with both detail features and spatial position information. After passing through the feature aggregation network, three feature maps F3, F4 and F5 are obtained.

7. The method for detecting underground targets based on multi-scale feature fusion according to claim 1, characterized in that: Step S4 specifically includes: S41, transform the feature map into and S42. Add 5 channels in the channel dimension, of which 4 channels are used to predict the position coordinates of the detection box, and 1 channel is used to determine whether there is an object at that position; S43, each pixel point spatial position passes through the multi-layer perceptron, and the number of channels is mapped to the probability of each category. The first 4 are [t x ,t y ,t w ,t h ], the mapping coordinates are calculated by the decoding formula, the decoding formula is as follows: x=s·(2σ(t x )-0.5+i) y=s·(2σ(t y )-0.5+j) w=s·(2σ(t w )) 2 h=s·(2σ(t h )) 2 Among them, s is the downsampling multiple, σ is the Sigmoid activation function, (i, j) is the feature map position of the prediction point, t x With t y Represents the position prediction value, (x, y) is the position of the prediction box in the input image, t w With t h are the predicted values ​​of width and height, respectively, w and h are the width and height of the predicted box in the input image; S44. The detection head performs detection box regression through coordinate mapping. This process uses a loss function to constrain it. The IoU loss function based on spatial perception is as follows: l WIoU =r WIoU ·(1-IoU) Among them, α is a positive balance parameter, v is used to measure the consistency of aspect ratio, and w gt With h gt is the width and height of the real detection frame, w and h are the width and height of the output detection frame, c is the diagonal length of the minimum circumscribed rectangle, b is the output detection frame position, b gt is the true detection box position, ρ(b,b gt ) represents the Euclidean distance between the center points, W g and H g is the width and height of the minimum bounding rectangle of the real detection box and the output detection box, in r WIoU In order to prevent the gradient from converging during the back propagation of the function, Separated from the computational graph and does not participate in back propagation; S45, after regression, multiple candidate detection frame sets B={B1,...,B i ,...B N }, where N represents the number of categories.

8. The method for detecting underground targets based on multi-scale feature fusion according to claim 1, characterized in that: Step S5 specifically includes: S51, according to the confidence threshold t, the i-th type detection box set B i Perform preliminary screening to obtain the detection frame group G i , G i ={g i1 ,...,g ij ,...,g iM }, where M represents the number of detection frames after preliminary screening, and initializes the output sequence D i , at this time D i is an empty set; S52, traverse the detection frame group G i , find the detection box M with the highest confidence, and add the detection box M to the output sequence D i ; S53, based on the detection frame M and the detection frame group G i The rest of the detection boxes g ij , calculate the intersection-over-union ratio, the calculation formula is as follows: Among them, IoU ij Represents the detection box M and the detection box group G i The j-th detection box g in ij The intersection and union ratio of S54, remove the detection frame M from the detection frame group G i Eliminate it and only compare each intersection over IoU ij The confidence of boxes greater than or equal to the threshold is linearly decayed, and the calculation formula is as follows: Among them, τ is the intersection-over-union ratio threshold, s ij is the confidence score of the j-th detection box of the i-th category, s′ ij is the confidence score after transformation; S55, from the detection frame group G i Remove the low confidence score s′ ij The corresponding detection frame, until the detection frame group G i There are no remaining candidate boxes, and the output D is obtained i ; S56: The remaining categories also perform the process from step S51 to step S55 to obtain the final output D={D1,...,D i ,...,D N }.

Citation Information

Cited By

  • A vehicle detection method, apparatus, equipment, and medium based on deformable convolution.

    CN122416385A