A small target visual detection method in geological exploration scenarios

By constructing a VDST-GA detection method for geological exploration scenarios, and utilizing generative adversarial networks and multimodal data fusion strategies, the complexity of small target detection in geological exploration is solved, achieving high-precision and efficient target recognition.

CN120612561BActive Publication Date: 2025-10-28CHENGDU UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511120661.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-10-28
Estimated Expiration
2045-08-12

AI Technical Summary

Technical Problem

Small target detection in geological exploration scenarios faces challenges such as interference from complex surface environments, blurred target features, diverse rock textures, and edge information degradation caused by weathering. Existing general target detection algorithms are also unable to adapt to the scale sensitivity of geological targets, leading to missed detections and false detections. Furthermore, the scarcity of labeled samples restricts the generalization ability of deep learning models.

Method used

A visual detection method for small targets in geological exploration scenarios, VDST-GA, is constructed, which includes a backbone network CSPDarknet, a geological target feature construction module, a feature fusion module ASFF, and a geological feature recognition module. The feature extraction and recognition capabilities are enhanced through generative adversarial networks and multimodal data fusion strategies.

Benefits of technology

It improves the accuracy of small target detection, reduces background noise interference, enhances the model's perception and adaptability to small target areas, and strengthens the automation and precision of geological exploration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612561B_ABST
    Figure CN120612561B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of image recognition technology and discloses a visual detection method for small targets in geological exploration scenarios. The method involves obtaining initial training samples containing various types of small targets to be detected in the geological field; constructing an initial deep learning network model, which includes a backbone network CSPDarknet, a hybrid feature fusion module ASFF, a geological target feature construction module, a geological feature recognition module, and a FastestDet detection head; inputting the initial training samples into the initial deep learning network model for training to obtain a trained deep learning network model; and inputting the image to be detected into the trained deep learning network model VDST-GA to output the predicted target location and target type. This invention overcomes the feature information loss in traditional target detection models when extracting small targets, making the model more focused on small target regions, reducing background noise interference, and offering advantages such as high accuracy, good adaptability, and high production efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and specifically to a visual detection method for small targets in geological exploration scenarios. Background Technology

[0002] With the acceleration of global industrialization and the surge in energy demand, geological exploration is playing an increasingly important role in mineral resource development, geological disaster early warning, and underground space utilization. Traditional geological exploration mainly relies on manual field surveys and geophysical equipment analysis, which has limitations such as low efficiency, high cost, and difficulty in covering high-risk areas. In recent years, intelligent exploration systems based on UAV remote sensing, mobile robot platforms, and intelligent sensing technologies have developed rapidly. By equipping these systems with high-resolution cameras, multispectral sensors, and other equipment, they can efficiently acquire high-precision image data of the surface and shallow geology, providing a data foundation for automated geological analysis.

[0003] However, key targets in geological scenes (such as mineral outcrops, microfractures, and dike boundaries) often exhibit characteristics such as small size (less than 1% of pixels), irregular shape, and low contrast with the background, making them typical small target detection problems. These targets play a crucial role in resource location and geological structure analysis, but their visual detection faces multiple challenges: 1) Complex surface environments (such as vegetation cover, uneven lighting, and shadow occlusion) blur target features; 2) Diverse rock textures and weathering degrade edge information of small targets; 3) Existing general-purpose target detection algorithms (such as Faster R-CNN and YOLO series) have feature pyramids and receptive field mechanisms designed for natural scenes that are difficult to adapt to the scale sensitivity of geological targets, easily leading to missed detections and false detections. Furthermore, the highly specialized nature of geological data results in a scarcity of labeled samples, further limiting the generalization ability of deep learning models.

[0004] To address the aforementioned issues, research on visual detection methods for small targets in geological exploration scenarios has significant application value. By constructing a domain-adaptive feature enhancement mechanism and a multimodal data fusion strategy, the recognition accuracy of key geological targets can be improved, providing core technical support for intelligent exploration equipment and playing a vital role in promoting the automation and precision transformation of geological surveys. Summary of the Invention

[0005] The purpose of this invention is to solve the problem of small target detection in geological exploration scenarios using existing deep learning network models, and to disclose a visual detection method for small targets in geological exploration scenarios, VDST-GA. VDST-GA comprises five parts: a backbone network CSPDarknet, a geological target feature construction module, a feature fusion module ASFF, a geological feature recognition module, and a FastestDet detection head. By inputting the image to be detected into the trained VDST-GA model, the VDST-GA model outputs the predicted small target detection location and target type. The geological target feature construction module is based on an attention-based generative adversarial network to compensate for the feature information lost during small target extraction; while the geological feature recognition module combines spatial attention and self-attention mechanisms to make the model more focused on the region and reduce background noise interference.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] A visual detection method for small targets in a geological exploration scenario includes the following steps:

[0008] S1, Obtain initial training samples, which contain small target types that need to be detected in the geological field;

[0009] S2, Construct the initial deep learning network model; the initial deep learning network model includes the backbone network CSPDarknet, the feature fusion module ASFF, the geological target feature construction module, the geological feature recognition module, and a FastestDet detection head;

[0010] The initial deep learning network model is constructed as follows:

[0011] S21, construct the backbone network CSPDarknet, the feature fusion module ASFF, and the FastestDet detection head;

[0012] S22, Construct the geological target feature construction module, which includes two sub-modules: a feature extraction module and a two-order kernel convolutional unit;

[0013] S23, Construct a geological feature recognition module, which combines the spectral information of the image and performs spatial attention, channel attention and self-attention on the spectral features of the image, and maps the multi-attention fused spectral features into image features;

[0014] S24, the geological target feature construction module, geological feature recognition module, feature fusion module ASFF and FastestDet detection head are embedded into the backbone network CSPDarknet module to obtain the initial deep learning network model;

[0015] S3, input the initial training sample set into the initial deep learning network model for training, and obtain the trained deep learning network model;

[0016] S4 inputs the image to be detected into the trained deep learning network model, which outputs the predicted location and type of the geological target.

[0017] In one possible implementation, in step S22, the geological target feature construction module is constructed based on a feature extraction module built using a generative network and a two-stage kernel convolutional unit. The two-stage kernel convolutional unit consists of a spatial kernel convolution and a point kernel convolution, specifically represented by the following formula:

[0018] (1);

[0019] In the formula, F Features representing the input two-stage kernel convolutional units; PC Represents a two-stage kernel convolutional unit; G Represents spatial kernel convolution; Z This represents point kernel convolution; This represents element-wise multiplication; furthermore, the number of parameters in a two-order kernel convolution unit... P DS and computational complexity C DS Represented as equations (2) and (3):

[0020] P DS =D W ×D H ×M+ 1 × 1 ×N×M (2);

[0021] C DS =(D W ×D H ×M+ 1 × 1 ×N×M)×F W ×F H (3);

[0022] in, F W and F H These represent the length and width of the feature map extracted by the feature extraction module, respectively. M The number of feature map channels extracted by the feature extraction module; NThis represents the number of convolution kernels; the size of the convolution kernel is... D W × D H × M ,in D W and D H These represent the length and width of the convolution kernel, respectively; 1×1× N × M represent N Each of size 1×1× M The convolution kernel is used for point-by-point convolution; thus, each channel is fused, and finally obtained. M Output feature map of each channel;

[0023] The formula (4) for calculating the adversarial loss function of the generative adversarial network is:

[0024] L = E x~Pr [ f w ( x )]- E x~Pg [ f w ( x (4)

[0025] in, f w ( x ) represents hyperparameters w The mean squared error function; E x~Pr Indicates from sampling point x~P r Obtain a sample x Expected value; E x~Pg Indicates from sampling point x~Pg Obtain a sample x Expected value; This represents the calculated adversarial loss value.

[0026] In one possible implementation, in step S22, the feature extraction module is constructed using an encoder-decoder structure. The feature extraction module is built based on a generative adversarial network, and three dynamic attention modules (DCBAMs) are designed to obtain weighted feature maps with different weights. F CBAM1 、F CBAM2 、F CBAM3The three attention DCBAM modules are located at the input, between the encoder and decoder, and at the output, respectively. Furthermore, to distinguish features in the data samples generated by the feature extraction module from those in the real data samples, a discriminator network is embedded at the end of the feature extraction module. The construction process of the feature extraction module is as follows:

[0027] S221, the input feature map is fed into the first DCBAM module to obtain the feature map after initial attention weighting. F CBAM1 , F∈R H×W×C The feature map is represented by a matrix. R And the size of the matrix is ​​( H , W , C ), H and W These represent the length and width of the matrix, respectively. C The number of channels in the matrix is ​​represented by the initial attention weighting, as shown in equation (5):

[0028] (5);

[0029] In the above formula, Channel attention1 This indicates the attention weight of the first channel; Spatial attention1 This represents the first spatial attention weight; F CBAM1 This represents the feature map after initial attention weighting; Indicates cumulative operation; This represents a feature-weighted operation; Dynamic (·) represents the gating decision function, and its specific calculation formula is shown in the following formula:

[0030] (6);

[0031] In the formula, F input Represents the input feature map; W 1. W 2 represents a learnable parameter; μ ( F input ) indicates to F input Perform channel average activation; σ ( F input )express F input Standard deviation; e(F input ) express Finput Shannon entropy; b 1 and b 2 represents the threshold for controlling the activation of the hidden layer;

[0032] S222, will F CBAM1 The data is fed into the encoder, which contains cubic 2D convolutions; after passing through the cubic 2D convolutions contained in the encoder, F CBAM1 The size is reduced to 0.5 times its original size, and the calculation formula is shown in equation (7):

[0033] (7);

[0034] In the above formula, Encoder Indicates encoder; F CBAM1 This represents the feature map after initial attention weighting; This represents the feature map after being re-encoded by the encoder.

[0035] S223, after obtaining the encoder's output... Then, it is fed into the second DCBAM module, where it is attention-weighted again to obtain a second attention-weighted feature map. F CBAM2 The calculation formula is shown in equation (8):

[0036] (8);

[0037] In the above formula, Dynamic(·) represents the gating decision function; Channel attention2 This represents the weight of the attention for the second channel; Indicates to Perform channel attention weighting; Spatial attention2 Indicates the weights of the second spatial attention; Indicates to Perform spatial attention weighting; Indicates cumulative operation; This represents a feature-weighted operation;

[0038] S224, will F CBAM2 The data is fed into the decoder. The decoder part contains a scale-invariant two-dimensional convolution and three upsampling operations. After the convolution in the decoder, the size of the feature becomes four times the original size. The calculation formula is shown in equation (9):

[0039] (9);

[0040] In the above formula, FCBAM2 This represents the feature map after a second attention weighting. Decoder Indicates decoder, This represents the feature map after decoding by the decoder;

[0041] S225, after obtaining Then, after attention-weighted processing by the third DCBAM module at the output, the feature map that finally recovers the feature information of the small target is obtained. F CBAM3 The calculation formula is shown in equation (10):

[0042] (10);

[0043] In the formula, Dynamic(·) represents the gating decision function, and Channel... attention3 The weights representing the attention of the third channel. Indicates to Perform channel attention weighting; Spatial attention3 The weights represent the third spatial attention. Indicates to Perform spatial attention weighting; Indicates cumulative operation. This represents a feature-weighted operation;

[0044] S226, In order to distinguish between the features extracted by the feature extraction module and the features in the real data samples, a discriminator network is also embedded at the end of the feature extraction module. The discriminator network includes a two-dimensional convolution and three downsampling operations, specifically: the features generated by the feature extraction module are processed by the discriminator network. F fake and real features F real The two vectors are fed into the discriminator network respectively. After obtaining the output vectors, the discriminant loss between them is calculated. The calculation formula is shown in Equation (11):

[0045] (11);

[0046] In the above formula, F fake This represents the features generated by the feature extraction module; F real || represents true features; F real and F fake Perform an OR operation between them; JUDG This represents the discriminator network; D output This represents the discrimination loss of the discriminator network's output.

[0047] Furthermore, BatchNorm2d normalization and LeakyReLU activation functions are added to the convolution operations in steps S222 and S224. By adding BatchNorm2d normalization in steps S222 and S224, this invention helps reduce internal covariate shifts, resulting in faster network convergence. It also helps alleviate gradient vanishing and exploding problems, thus improving gradient propagation. The LeakyReLU activation function allows for a small gradient in the negative range, which not only helps alleviate the gradient vanishing problem but also makes it easier for the generator to learn the distribution of features and generate more diverse samples. Therefore, adding BatchNorm2d and LeakyReLU activation functions makes the generator training more stable. The use of the Tanh activation function after the last convolution operation in step S224 aims to accelerate the generator's convergence.

[0048] Furthermore, in step S226, the 2D convolution operation is augmented with InstanceNorm2d normalization and the LeakyReLU activation function. In the discriminator network, InstanceNorm2d normalization replaces the BatchNorm2d normalization used in the feature extraction module. This is because in generative adversarial networks, the discriminator needs to learn to distinguish subtle differences between real and generated samples. Using InstanceNorm2d better preserves the individual features of each sample without mixing them with other statistical information, which helps improve the discriminator's pattern recognition ability. Moreover, compared to BatchNorm2d normalization, InstanceNorm2d normalizes on a single sample, making it more suitable for use in the discriminator and helping to reduce the risk of overfitting.

[0049] Furthermore, in step S22, the feature extraction module needs to be trained; the entire training process is divided into three stages:

[0050] In the first stage, the baseline model CSPDarknet was trained using the initial training samples.

[0051] In the second stage, the original training sample images and the images reduced by 1 / 2 are respectively fed into the benchmark model CSPDarknet; the feature map A obtained after feeding the original image into the benchmark model CSPDarknet is used as the supervision signal for training the feature extraction module, i.e., the real feature in formula (11); the feature map B obtained after feeding the image reduced by 1 / 2 into the benchmark model CSPDarknet is used as the input of the feature extraction module.

[0052] In the third stage, feature maps are generated based on the feature extraction module, and the discriminant loss is obtained through formula (11). The feature extraction module is then optimized based on the discriminant loss. At the same time, the adversarial loss is obtained through formula (4), and the discriminator network is then optimized based on the adversarial loss.

[0053] This invention cross-trains the feature extraction module and discriminator of the entire generative adversarial network to achieve an optimal state for each. This training method enables the resulting feature extraction module to extract minute feature information.

[0054] Furthermore, the geological target feature construction module also includes an output unit, which combines the outputs of the feature extraction module and the two-stage kernel convolution unit with the output of the feature fusion module ASFF and performs a two-dimensional convolution operation.

[0055] In one possible implementation, in step S23, the construction process of the geological feature identification module is as follows:

[0056] S231, firstly, the spectral data of the feature map is calculated, and spatial attention is used to capture the spatial dependencies between different locations in the spectrum, simultaneously focusing on the characteristics of different regions across the entire image space; specifically, the input feature map... F∈R H×W×C Different spectral decomposition spatial attention methods are applied to focus on small target regions from different dimensions and perceive the characteristics of these regions. R H×W×C The feature map is represented by a matrix. R And the size of the matrix is ​​( H , W , C ), H and W These represent the length and width of the matrix, respectively. C The number of channels representing the matrix is ​​calculated using the formula shown in equation (12):

[0057] F spatial_n = SpAttention ( G sw ( F ))(12);

[0058] in, G sw Representative of feature maps F Calculate its spectral data to achieve the mapping from image features to spectral features; F spatial_n Represents the spatial attention spectral feature map, where the subscripts are... n This represents the dimension of spatial attention; SpAttention Represents the spectral feature map of the input.G sw ( F Spatial attention is performed in the corresponding different dimensions;

[0059] S232, for feature maps F∈R H×W×C After calculating its spectral data, channel attention is applied to each channel of the spectral feature matrix to obtain the channel attention spectral feature map. F ch The calculation formula is shown in equation (13):

[0060] F ch =Ch_Attention ( G sw ( F ))(13;

[0061] in, Ch_Attention Indicates channel attention; G sw ( F () represents a spectral feature map; F ch This represents the channel attention spectral feature map;

[0062] S233, for feature maps F∈R H×W×C After calculating its spectral data, self-attention is applied to all channels of the spectral feature matrix to obtain the self-attention spectral feature map. F self The calculation formula is shown in equation (14):

[0063] F self =Self_Attention ( G sw ( F ))(14;

[0064] in, Self Attention Indicates self-attention; G sw ( F () represents a spectral feature map; F self Represents the self-attention spectral feature map;

[0065] S234, perform spectral adaptive fusion on the above spatial attention spectral feature map, channel attention spectral feature map, and self-attention spectral feature map, and the calculation formula is shown in the following formula:

[0066] (15);

[0067] In the formula, F O This indicates the final feature obtained; F k ∈{ F spatial_n , F ch , F self}; MAP ( F k ) indicates that the spectral feature map F k Mapping spectral features to image features; W k , σ k , μ k This represents the weights, variances, and means of the corresponding feature maps, i.e., when... F k When representing spatial attention spectral features, W k Represents spatial attention weights. σ k The variance representing the spectral characteristics of spatial attention. μ k The mean value representing the spectral characteristics of spatial attention; ∑ indicates Euclidean operations; ∑ indicates accumulation operations.

[0068] In one possible implementation, in step S24, the geological target feature construction module, the geological feature recognition module, and the feature fusion module ASFF are embedded into the backbone network CSPDarknet module and combined with a FastestDet detector head to obtain an initial deep learning network model. The initial deep learning network model consists of five parts: the backbone network CSPDarknet, the feature fusion module ASFF, the FastestDet detector head, the geological target feature construction module, and the geological feature recognition module. The backbone network CSPDarknet includes six layers arranged sequentially. The initial deep learning network model construction process is as follows:

[0069] S241, the output feature map of the fourth layer of the backbone network CSPDarknet is used as the input of the geological feature recognition module, and then the output of the geological feature recognition module and the output of the fourth layer are used as the input of the fifth layer of the backbone network CSPDarknet.

[0070] S242, the outputs of the fourth, fifth, and sixth layers of the backbone network CSPDarknet are used as input features for the feature fusion module ASFF. Then, the ASFF module performs fusion processing on the input features and outputs a feature map. The feature fusion module ASFF includes three layers of output feature maps with progressively smaller scales: fusion. Figure 1 Integration Figure 2 and integration Figure 3 ;

[0071] S243, the input to the geological target feature construction module is divided into two parts. The first part is the output feature map of the fourth layer of the backbone network CSPDarknet; the second part is the fusion of the feature fusion module ASFF. Figure 1 The first part of the input serves as the input to the feature extraction module and the double-kernel convolution unit; the second part of the input, along with the outputs of the feature extraction module and the double-kernel convolution unit, serves as the input to the output unit; the output of the output unit serves as the output of the geological target feature construction module.

[0072] S244, the output of the feature fusion module ASFF and the geological target feature construction module, serves as the input of the FastestDet detection head; the FastestDet detection head outputs the category and location information of the target in the image.

[0073] Furthermore, the backbone network CSPDarknet has the following structure: the first layer is a focus layer; the second layer consists of two-dimensional convolution, normalization, and activation functions; the third layer consists of two-dimensional convolution, normalization, activation functions, and convolutional layers; the fourth layer consists of two-dimensional convolution, normalization, activation functions, and convolutional layers; the fifth layer consists of two-dimensional convolution, normalization, activation functions, and convolutional layers; and the sixth layer consists of two-dimensional convolution, normalization, activation functions, spatial pyramid pooling, and convolutional layers.

[0074] Further, in step S3, the initial deep learning network model needs to be trained: First, load and lock the training parameters of the baseline model CSPDarknet and the feature extraction module that have been trained in step 23; then, input the initial training samples into the initial deep learning network model, obtain the prediction results (small target location and target type) according to the steps S241-S244 given above, obtain the loss value of the initial deep learning network model based on the prediction results, and then optimize the parameters of the initial deep learning network model based on the loss value.

[0075] Furthermore, in step S3, the total loss function used by the initial deep learning network model VDST-GA is the loss function of the disclosed YoLo model, specifically the cross-entropy loss function.

[0076] Compared with the prior art, the present invention has the following beneficial effects:

[0077] 1. The small target visual detection method proposed in this invention addresses the issue of small target size in some geological exploration scenarios by incorporating an attention mechanism into a generative adversarial network to form a feature extraction module. This allows the trained feature extraction module to more accurately capture the feature details of tiny targets when extracting features of small targets. Furthermore, by adding the trained feature extraction module to the geological target feature construction module, feature information of small targets that are lost during downsampling can be extracted, thereby improving the accuracy of small target detection.

[0078] 2. The small target visual detection method proposed in this invention addresses the problem of insufficient perception of small target areas caused by the small size of some targets in certain geological exploration scenarios. It constructs a geological feature recognition module based on channel attention, spatial attention, and self-attention mechanisms. Unlike existing technologies where a single spatial attention may only focus on important target areas from a single dimension, this invention can provide richer contextual information, enabling the model to pay detailed attention to the local features of small targets while making full use of global information, thereby improving the perception ability of small target areas and playing a very important role in further improving target accuracy.

[0079] 3. The small target visual detection method proposed in this invention embeds a geological target feature construction module, a geological feature recognition module, and a feature fusion module ASFF into the backbone network CSPDarknet. The constructed deep learning network model makes up for the feature information lost by traditional target detection models in extracting small targets, makes the model pay more attention to small target areas, reduces the interference of background noise, and has the advantages of high accuracy, good adaptability and high production efficiency. Attached Figure Description

[0080] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0081] Figure 1 This is a flowchart of the small target visual detection method in geological exploration scenarios proposed in this invention;

[0082] Figure 2 Structure diagram of the module for constructing geological target features;

[0083] Figure 3 Here is a structural diagram of the feature extraction module;

[0084] Figure 4A structural diagram of the geological target feature recognition module;

[0085] Figure 5 This is a diagram of the structure of a deep learning network model. Detailed Implementation

[0086] In the following description, only certain exemplary embodiments are briefly described. As those skilled in the art will recognize, the described embodiments can be modified in various ways without departing from the spirit or scope of the invention. Therefore, the drawings and description are considered to be exemplary in nature and not restrictive.

[0087] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0088] This invention provides a visual detection method for small targets in geological exploration scenarios, such as... Figure 1 As shown, it includes the following steps:

[0089] S1, Obtain the initial training samples.

[0090] 2000 small target images of geological exploration scenes were taken with a camera, and the images were divided into training sample set and test sample set according to a 7:3 ratio.

[0091] S2, Construct the initial deep learning network model.

[0092] The initial deep learning network model VDST-GA includes a backbone network CSPDarknet, a feature fusion module ASFF, a FastestDet detector head, a geological target feature construction module, and a geological feature recognition module. The feature fusion module ASFF, the geological target feature construction module, and the geological feature recognition module are embedded within the backbone network CSPDarknet. The resulting deep learning network model is shown below. Figure 5 As shown.

[0093] Step S2 includes the following sub-steps:

[0094] S21, construct the backbone network CSPDarknet, the feature fusion module ASFF, and the FastestDet detection head.

[0095] like Figure 5As shown, the backbone network CSPDarknet consists of six layers arranged sequentially: the first layer is a focus layer; the second layer is a 2D convolutional layer, normalization layer, and activation function layer; the third layer is a 2D convolutional layer, normalization layer, activation function layer, and convolutional layer; the fourth layer is a 2D convolutional layer, normalization layer, activation function layer, and convolutional layer; the fifth layer is a 2D convolutional layer, normalization layer, activation function layer, and convolutional layer; and the sixth layer is a 2D convolutional layer, normalization layer, activation function layer, spatial pyramid pooling layer, and convolutional layer. The convolutional layers are 1D, and the activation function used is the SiLU activation function. The outputs of the fourth, fifth, and sixth layers of the backbone network CSPDarknet serve as the three outputs of the backbone network CSPDarknet, namely feature maps feat1, feat2, and feat3, respectively.

[0096] The feature fusion module ASFF adopts the existing classical feature fusion method ASFF (see: Al-Haddad AA, Ettouney HM, Abu-Irhayem TM. Modeling of Degradation of Solutions Containing Nitrogen and Carbon Compounds in Aerated Submerged Fixed-Film (ASFF) Bioreactors[J]. Environmental Technology Letters, 1996, 17(2):113-124.DOI:10.1080 / 09593331708616368); the feature fusion module ASFF includes three layers of output with gradually decreasing scale: fusion Figure 1 Integration Figure 2 and integration Figure 3 .

[0097] The FastestDet detector head uses a Yolox neural network to predict the type and location of the target being detected, based on existing detector heads.

[0098] In step S22, the geological target feature construction module is constructed based on the feature extraction module built by the generative network and the two-stage kernel convolution unit. The two-stage kernel convolution unit consists of a spatial kernel convolution and a point kernel convolution, which can be specifically represented by the following formula:

[0099] (1);

[0100] In the formula, F The feature map feat1 represents the features of the input two-stage kernel convolutional unit. PC This represents a two-stage kernel convolutional unit. GRepresents spatial kernel convolution; Z This represents point kernel convolution; This represents element-wise multiplication; furthermore, the number of parameters in a two-order kernel convolution unit... P DS and computational complexity C DS Represented as equations (2) and (3):

[0101] P DS =D W ×D H ×M+ 1 × 1 ×N×M (2);

[0102] C DS =(D W ×D H ×M+ 1 × 1 ×N×M)×F W ×F H (3);

[0103] in, F W and F H These represent the length and width of the feature map extracted by the feature extraction module, respectively. M The number of feature map channels extracted by the feature extraction module; N This represents the number of convolution kernels; the size of the convolution kernel is... D W × D H × M ,in D W and D H These represent the length and width of the convolution kernel, respectively; 1×1× N × M represent N Each of size 1×1× M The convolution kernel is used for point-by-point convolution; thus, each channel is fused, and finally obtained. M Output feature map of each channel;

[0104] The formula (4) for calculating the adversarial loss function of the generative adversarial network is:

[0105] L=E x~Pr [ fw ( x )] -E x~Pg [ f w ( x (4)

[0106] in, f w (x) Represents hyperparameters w The mean squared error function; E x~Pr Indicates from sampling point x~P r Obtain a sample x Expected value; E x~Pg Indicates from sampling point x~Pg Obtain a sample x Expected value; This represents the calculated adversarial loss value.

[0107] In step S22, the feature extraction module is constructed using an encoder-decoder structure. The feature extraction module is built based on a generative adversarial network (GAN) and uses three dynamic attention modules (DCBAM) to obtain weighted feature maps with different weights. F CBAM1 、F CBAM2 、F CBAM3 The three attention DCBAM modules are located at the input end, between the encoder and decoder, and at the output end, respectively. In addition, in order to distinguish the features in the data samples generated by the feature extraction module from the features in the real data samples, a discriminator network is also embedded at the end of the feature extraction module.

[0108] The feature extraction module structure diagram is as follows: Figure 3 As shown, its construction process is as follows:

[0109] S221, the input feature map is fed into the first DCBAM module to obtain the feature map after initial attention weighting. F CBAM1 , F∈R H×W×C The feature map is represented by a matrix. R And the size of the matrix is ​​( H , W , C ), H and W These represent the length and width of the matrix, respectively. C The number of channels in the matrix is ​​represented by the initial attention weighting, as shown in equation (5):

[0110] (5);

[0111] In the above formula, Channel attention1 This indicates the attention weight of the first channel; Spatial attention1 This represents the first spatial attention weight; F CBAM1 This represents the feature map after initial attention weighting; Indicates cumulative operation; This represents a feature-weighted operation; Dynamic (·) represents the gating decision function, and its specific calculation formula is shown in the following formula:

[0112] (6);

[0113] In the formula, F input Represents the input feature map; W 1. W 2 is a learnable parameter. μ ( F input ) indicates to F input Perform channel average activation. σ ( F input )express F input Standard deviation; e(F input ) express F input Shannon entropy; b 1 and b 2 represents the threshold for controlling the activation of the hidden layer;

[0114] S222, will F CBAM1 The data is fed into the encoder, which contains cubic 2D convolutions; after passing through the cubic 2D convolutions contained in the encoder, F CBAM1 The size is reduced to 0.5 times its original size, and the calculation formula is shown in equation (7):

[0115] (7);

[0116] In the above formula, Encoder Indicates encoder; F CBAM1 This represents the feature map after initial attention weighting; This represents the feature map after being re-encoded by the encoder.

[0117] S223, after obtaining the encoder's output... Then, it is fed into the second DCBAM module, where it is attention-weighted again to obtain a second attention-weighted feature map. F CBAM2 The calculation formula is shown in equation (8):

[0118] (8);

[0119] In the above formula, Dynamic (·) represents the gating decision function; Channel attention2 This represents the weight of the attention for the second channel; Indicates to Perform channel attention weighting; Spatial attention2 Indicates the weights of the second spatial attention; Indicates to Perform spatial attention weighting; Indicates cumulative operation; This represents a feature-weighted operation;

[0120] S224, will F CBAM2 The data is fed into the decoder. The decoder part contains a scale-invariant two-dimensional convolution and three upsampling operations. After the convolution in the decoder, the size of the feature becomes four times the original size. The calculation formula is shown in equation (9):

[0121] (9);

[0122] In the above formula, F CBAM2 This represents the feature map after a second attention weighting. Decoder Indicates decoder; This represents the feature map after decoding by the decoder;

[0123] S225, after obtaining Then, after attention weighting by DCBAM at the output end, the feature map that finally recovers the feature information of the small target is obtained. F CBAM3 The calculation formula is shown in equation (10):

[0124] (10);

[0125] In the formula, Dynamic (·) represents the gating decision function; Channel attention3 The weights representing the attention of the third channel. Indicates to Perform channel attention weighting; Spatial attention3 The weights represent the third spatial attention. Indicates to Perform spatial attention weighting; Indicates cumulative operation; This represents a feature-weighted operation;

[0126] S226, In order to distinguish between the features extracted by the feature extraction module and the features in the real data samples, a discriminator network is also embedded at the end of the feature extraction module. The discriminator network includes a two-dimensional convolution and three downsampling operations, specifically: the features generated by the feature extraction module are processed by the discriminator network. F fake and real features F real The two vectors are fed into the discriminator network respectively. After obtaining the output vectors, the discriminant loss between them is calculated. The calculation formula is shown in Equation (11):

[0127] (11);

[0128] In the above formula, F fake This represents the features generated by the feature extraction module; F real || represents true features; F real and F fake Perform an OR operation between them; JUDG This represents the discriminator network; D output This represents the discrimination loss of the discriminator network's output.

[0129] Furthermore, BatchNorm2d normalization and LeakyReLU activation functions are added to the convolution operations in steps S222 and S224. By adding BatchNorm2d normalization in steps S222 and S224, this invention helps reduce internal covariate shifts, resulting in faster network convergence. It also helps alleviate gradient vanishing and exploding problems, thus improving gradient propagation. The LeakyReLU activation function allows for a small gradient in the negative range, which not only helps alleviate the gradient vanishing problem but also makes it easier for the generator to learn the distribution of features and generate more diverse samples. Therefore, adding BatchNorm2d and LeakyReLU activation functions makes the generator training more stable. The use of the Tanh activation function after the last convolution operation in step S224 aims to accelerate the generator's convergence.

[0130] Furthermore, in step S226, the 2D convolution operation is augmented with InstanceNorm2d normalization and the LeakyReLU activation function. In the discriminator network, InstanceNorm2d normalization replaces the BatchNorm2d normalization used in the feature extraction module. This is because in generative adversarial networks, the discriminator needs to learn to distinguish subtle differences between real and generated samples. Using InstanceNorm2d better preserves the individual features of each sample without mixing them with other statistical information, which helps improve the discriminator's pattern recognition ability. Moreover, compared to BatchNorm2d normalization, InstanceNorm2d normalizes on a single sample, making it more suitable for use in the discriminator and helping to reduce the risk of overfitting.

[0131] In this embodiment, the feature extraction module needs to be trained to obtain the ability to recover the feature information of small targets.

[0132] The entire training process is divided into three stages:

[0133] In the first stage, the baseline model CSPDarknet is trained using initial training samples until the set number of iterations is reached. After each iteration, the loss value is obtained through the loss function, and the baseline model CSPDarknet is optimized using the gradient descent optimization algorithm. The baseline model CSPDarknet uses the existing combined loss function of the traditional YoLo model (see Redmon J, Divvala S, Girshick R, et al. Youonlylookonce: Unified, real-time object detection [C]. Proceedings of the IEEE conference on computer vision and pattern recognition. 2016:779-788).

[0134] In the second stage, the original images and the images reduced by 1 / 2 of the initial training samples are fed into the benchmark model CSPDarknet respectively; the feature map A obtained after feeding the original images into the benchmark model CSPDarknet is used as the supervision signal for training the attention generative adversarial network, i.e., the real features in formula (11); the feature map B obtained after feeding the images reduced by 1 / 2 into the benchmark model CSPDarknet is used as the input of the feature extraction module.

[0135] In the third stage, features are generated based on the feature extraction module, and the discriminant loss is obtained through formula (11). The feature extraction module is then optimized based on the discriminant loss. At the same time, the adversarial loss is obtained through formula (4), and the discriminator network is then optimized based on the adversarial loss.

[0136] Repeat the second and third stages as described above until the set number of iterations is reached.

[0137] This invention cross-trains the feature extraction module and discriminator of the entire generative adversarial network to achieve an optimal state for each. This training method enables the resulting feature extraction module to extract minute feature information.

[0138] like Figure 2 As shown, the trained feature extraction module is integrated with a dual-kernel convolutional unit to form a geological target feature construction module. The geological target feature construction module also includes an output unit, which combines the outputs of the feature extraction module and the dual-kernel convolutional unit with the output of the feature fusion module ASFF (i.e., fusion). Figure 1 The feature extraction module and the two-dimensional convolution unit both take the feature map feat1 as input.

[0139] S23, construct a geological feature recognition module, which combines the spectral information of the image and applies spatial attention, channel attention, and self-attention to the spectral features of the image, and then maps the multi-attention fused spectral features into image features; such as Figure 4 As shown, the construction process of the geological feature recognition module is as follows:

[0140] S231, firstly, the spectral data of the feature map is calculated, and spatial attention is used to capture the spatial dependencies between different locations in the spectrum, simultaneously focusing on the characteristics of different regions across the entire image space; specifically, the input feature map... F∈R H×W×C Different spectral decomposition spatial attention methods are applied to focus on small target regions from different dimensions and perceive the characteristics of these regions. R H×W×C The feature map is represented by a matrix. R And the size of the matrix is ​​( H , W , C ), H and W These represent the length and width of the matrix, respectively. C The number of channels representing the matrix is ​​calculated using the formula shown in equation (12):

[0141] F spatial_n = SpAttention ( G sw( F ))(12);

[0142] in, G sw Representative of feature maps F Calculate its spectral data to achieve the mapping from image features to spectral features; F spatial_n Represents the spatial attention spectral feature map, where F spatial_n subscript n This represents the dimension of spatial attention; SpAttention Represents the spectral feature map of the input. G sw ( F Spatial attention is performed in the corresponding different dimensions;

[0143] S232, for feature maps F∈R H×W×C After calculating its spectral data, channel attention is applied to each channel of the spectral feature matrix to obtain the channel attention spectral feature map. F ch The calculation formula is shown in equation (13):

[0144] F ch =Ch_Attention ( G sw ( F ))(13;

[0145] in, Ch_Attention Indicates channel attention; G sw ( F () represents a spectral feature map; F ch This represents the channel attention spectral feature map;

[0146] S233, for feature maps F∈R H×W×C After calculating its spectral data, self-attention is applied to all channels of the spectral feature matrix to obtain the self-attention spectral feature map. F self The calculation formula is shown in equation (14):

[0147] F self =Self_Attention ( G sw ( F ))(14;

[0148] in, Self Attention Indicates self-attention; G sw ( F () represents a spectral feature map; F self Represents the self-attention spectral feature map;

[0149] S234, perform spectral adaptive fusion on the above spatial attention spectral feature map, channel attention spectral feature map, and self-attention spectral feature map, and map the fusion result from spectral features to image features. The calculation formula is as follows:

[0150] (15);

[0151] In the formula, F O This indicates the final feature obtained; F k ∈{ F spatial_n , F ch , F self}; MAP ( F k ) indicates that the spectral feature map F k Mapping spectral features to image features; W k , σ k , μ k This represents the weights, variances, and means of the corresponding feature maps, i.e., when... F k When representing spatial attention spectral features, W k Represents spatial attention weights. σ k The variance representing the spectral characteristics of spatial attention. μ k The mean value representing the spectral characteristics of spatial attention; ∑ indicates Euclidean operations; ∑ indicates accumulation operations.

[0152] The above features map F Calculate its spectral data G swThe conventional methods disclosed in this field were adopted. For specific operations, please refer to Zhou H, Shen Y, Yan Z, et al. Generative adversarial network with double discrimination for heterogeneous hyperspectral reconstruction[J]. Computers and Electronics in Agriculture, 237[2025-07-27].

[0153] The above will show the spectral feature map F k The mapping from spectral features to image features (MAP) employs conventional methods already disclosed in the field. For details, please refer to Liao J, Wang L. HyperspectralMamba: A Novel State SpaceModel Architecture for Hyperspectral Image Classification[J]. Remote Sensing,2025, 17(15): 2577.

[0154] The initial deep learning network model VDST-GA construction process is as follows:

[0155] S241, the output feature map of the fourth layer of the backbone network CSPDarknet is used as the input of the geological feature recognition module, and then the output of the geological feature recognition module and the output of the fourth layer are used as the input of the fifth layer of the backbone network CSPDarknet.

[0156] S242, the outputs of the fourth, fifth, and sixth layers of the backbone network CSPDarknet are used as input features for the feature fusion module ASSF. Then, the feature fusion module ASSF performs fusion processing on the input features and outputs a feature map. The feature fusion module ASSF includes three layers of output feature maps with progressively smaller scales: fusion. Figure 1 Integration Figure 2 and integration Figure 3 ;

[0157] S243, the input to the geological target feature construction module is divided into two parts. The first part is the output feature map of the fourth layer of the backbone network CSPDarknet; the second part is the fusion of the feature fusion module ASFF. Figure 1The first part of the input serves as the input to the feature extraction module and the double-kernel convolution unit; the second part of the input, along with the outputs of the feature extraction module and the double-kernel convolution unit, serves as the input to the output unit; the output of the output unit serves as the output of the geological target feature construction module.

[0158] S244, the output of the feature fusion module ASSF and the geological target feature construction module, serves as the input of the FastestDet detection head; the FastestDet detection head outputs the category and location information of the target in the image.

[0159] After feature extraction by the backbone network CSPDarkNet, feature fusion is performed using the enhanced feature fusion module ASFF, such as... Figure 5 As shown, when the input image size is (640, 640, 3), CSPDarkNet will output three feature maps of different sizes: feat1, feat2, and feat3. These feature maps are located in the middle layer, the lower middle layer, and the bottom layer, respectively. For example, the sizes of the three feature maps are feat1=(80, 80, 256), feat2=(40, 40, 512), and feat3=(20, 20, 1024). Through the feature fusion module ASFF, the initial deep learning network model can effectively fuse these feature maps of different sizes to obtain rich feature representations. For details of the feature fusion process, please refer to Al-Haddad AA, Ettouney HM, Abu-Irhayem TM. Modeling of Degradation of Solutions Containing Nitrogen and Carbon Compounds in Aerated Submerged Fixed-Film (ASFF) Bioreactors[J].Environmental Technology Letters, 1996, 17(2):113-124.DOI:10.1080 / 09593331708616368.

[0160] The initial deep learning network model structure constructed in this embodiment is as follows: Figure 5 As shown, a geological target feature construction module, a geological feature recognition module, and a feature fusion module ASFF are embedded in the backbone network CSPDarknet. This invention chooses CSPDarknet as the backbone network because it has four important characteristics:

[0161] (1) A residual network, Residual, is used, where the residual convolution can be split into two parts: the backbone consists of 1×1 convolutions and 3×3 convolutions. The residual part remains unchanged, directly combining the input and output of the backbone network. In deep neural networks, as the number of network layers increases, the problems of vanishing and exploding gradients become more and more serious. Residual networks alleviate the problems of vanishing and exploding gradients by introducing skip connections, allowing information to be directly passed from one layer to another, thus making the network easier to train.

[0162] (2) The CSPnet network structure was used to re-divide the original residual block stack into two parts: the main part continues to stack the original residual blocks, while the other part is similar to a residual edge, which is directly connected to the final output after a small amount of processing.

[0163] (3) A Focus network structure is used, which can compress and reconstruct the input feature map. Specifically, the input feature map is first downsampled to 0.5 of its original size and then divided into four sub-feature maps. This downsampling process can be implemented using ordinary convolution operations. Then, each sub-feature map is assigned to a different convolution kernel for processing, and each convolution kernel only processes the channel dimension of the sub-feature map. Finally, the results of each sub-feature map are concatenated to obtain the final output features. The main purpose of this structure is to reduce computation and memory usage while maintaining model performance. Through downsampling and channel segmentation, the model can process larger input images with less space and time overhead without losing too much information. In addition, the number of convolution kernels assigned to each sub-feature map is relatively small, which can effectively reduce computational complexity and improve model speed.

[0164] (4) The SiLU activation function is used. The SiLU activation function is a combination of the Sigmoid activation function and the ReLU activation function. The SiLU function has the characteristics of smoothness and non-monotonicity. In CNN, the overall performance of the SiLU activation function is better than that of the Sigmoid activation function and the ReLU activation function.

[0165] S3. Input the initial training sample set into the initial deep learning network model VDST-GA for training to obtain the trained deep learning network model VDST-GA.

[0166] First, load and lock the training parameters of the baseline model CSPDarknet and the feature extraction module that have been trained in step 23.

[0167] Then, the initial training samples are input into the initial deep learning network model VDST-GA. Following steps S241-S244 given earlier, prediction results (target location and target type) are obtained. Based on the prediction results, the loss value of the initial deep learning network model VDST-GA is obtained. Then, based on the loss value, the parameters of the initial deep learning network model VDST-GA are optimized using the gradient descent optimization algorithm. This mainly optimizes the training parameters of network structures other than the baseline model CSPDarknet and the feature extraction module. This operation is repeated until the set number of iterations is reached.

[0168] Since this step mainly focuses on training the model's classification ability, the loss function used in the initial deep learning network model VDST-GA is the loss function of the already disclosed YoLo model, specifically the cross-entropy loss function.

[0169] S4. Input the image to be detected into the trained deep learning network model VDST-GA, which outputs the predicted location and type of the small target.

[0170] The image is input into the deep learning network model, and the model outputs the predicted location and type.

[0171] To demonstrate the effectiveness of this invention, the proposed VDST-GA model was compared with several of the most popular methods on the NWPUVHR-10 public dataset. The results are shown in Table 1.

[0172] Table 1. Comparative experimental results (%) on the NWPUVHR-10 dataset

[0173]

[0174] As shown in Table 1, compared with other classic small object detection models, this invention performs better on the NWPUVHR-10 dataset. AP small and AP 0.5 The VDST-GA achieved accuracy rates of 29.8% and 79.3% respectively, surpassing all other algorithms listed in Table 1. Compared to the best-performing YOLOv7-tiny model among other algorithms, VDST-GA... AP small An increase of 6.1%, AR small It increased by 8.9%. AP 0.5 It increased by 2.5%.

[0175] Compared to the NanoDet-m and NanoDet-m-1.5x models, VDST-GA performs better. AP smalland AP 0.5 It significantly outperforms both of these algorithms. Meanwhile, in AP , AP 0.75 , AP medium and AP large These four metrics also significantly outperformed the NanoDet-m and NanoDet-m-1.5x models. However, in AR small , AR medium , AR large In these three metrics, VDST-GA lags behind the NanoDet-m and NanoDet-m-1.5x models. This is primarily because both NanoDet-m and NanoDet-m-1.5x models employ techniques such as grouped convolution, depthwise separable convolution, and channel rearrangement. These techniques significantly compress and abstract the input features, resulting in better local feature extraction. However, in this process, some detailed information and local features may be lost due to compression, leading to a decrease in detection accuracy.

[0176] Compared to the NanoDet-EfficientNet-Lite1 model, VDST-GA performs better. AP small and AP 0.5 These figures represent increases of 9.2% and 20.5% respectively. Meanwhile, in and AP 0.75 These two metrics significantly outperform the NanoDet-EfficientNet-Lite1 model. However, at medium and large scales... and In terms of metrics, it lags behind the NanoDet-EfficientNet-Lite1 model. This is primarily because the NanoDet-EfficientNet-Lite1 model uses the EfficientNetLite network as its backbone. EfficientNetLite networks typically employ deep structures, extracting abstract features from images through multiple convolutions and pooling. As the number of layers increases, the feature resolution gradually decreases. This leads to the loss of detailed information about small objects in high-level feature maps. Therefore, the NanoDet-EfficientNet-Lite1 model outperforms the NanoDet-EfficientNet-Lite1 model in detecting medium- and large-scale objects.

[0177] Compared to the latest YOLOv7-tiny and YOLOv8-nano, VDST-GA performs better in the YOLOv7-tiny category. AP small , AP 0.5 and AR small The improvements were 6.1%, 2.5%, and 9.1% respectively. Regarding YOLOv8-nano, VDST-GA showed improvements in... AP small , AP 0.5 and AR small The improvements were 7.2%, 8.0%, and 10.1%, respectively. Although VDST-GA lags behind YOLOv7-tiny and YOLOv8-nano in some other metrics, the differences are not significant. Considering real-time performance, the VDST-GA model has significantly fewer parameters and lower computational cost than the YOLOv7-tiny and YOLOv8-nano models. Therefore, taking all factors into account, it is more suitable for applications in geological scenarios.

[0178] This invention proposes a visual detection method for small targets in geological exploration scenarios, which mainly consists of a geological target feature construction module and a geological feature recognition module. First, a proposed feature extraction module is integrated into the geological target feature construction module to extract small target feature information lost by the model during downsampling. Simultaneously, to make the model focus more on small target regions in the image, a geological feature recognition module is designed by combining spectral spatial attention, spectral channel attention, and spectral self-attention, enabling the model to pay closer attention to small target regions and reduce background noise interference. Finally, to mitigate the additional computational and parameter requirements brought about by adding modules, a two-order kernel convolution method is employed, making it practically applicable to real-world geological scenarios.

Claims

1. A visual detection method for small targets in a geological exploration scenario, characterized in that: Includes the following steps: S1, Obtain the initial training samples; S2, Construct the initial deep learning network model; the initial deep learning model includes the backbone network CSPDarknet, the feature fusion module ASFF, the geological target feature construction module, the geological feature recognition module, and a FastestDet detection head; The initial deep learning network model is constructed as follows: S21, construct the backbone network CSPDarknet, the feature fusion module ASFF, and the FastestDet detection head; S22, Construct the geological target feature construction module, which includes two sub-modules: a feature extraction module and a two-order kernel convolutional unit; The feature extraction module is constructed using an encoder-decoder structure. It is based on a generative adversarial network and employs three dynamic attention modules (DCBAMs) to acquire weighted feature maps with different weights. F CBAM1 、F CBAM2 、F CBAM3 The three attention DCBAM modules are located at the input, between the encoder and decoder, and at the output, respectively. Furthermore, to distinguish features in the data samples generated by the feature extraction module from those in the real data samples, a discriminator network is embedded at the end of the feature extraction module. The construction process of the feature extraction module is as follows: S221, the input feature map is fed into the first DCBAM module to obtain the feature map after initial attention weighting. F CBAM1 , F∈R H×W×C The feature map is represented by a matrix. R And the size of the matrix is ​​( H , W , C ), H and W These represent the length and width of the matrix, respectively. C The number of channels in the matrix is ​​represented by the initial attention weighting, as shown in equation (5): (5); In the above formula, Channel attention1 This indicates the attention weight of the first channel; Spatial attention1 This represents the first spatial attention weight; F CBAM1 This represents the feature map after initial attention weighting; Indicates cumulative operation; This represents a feature-weighted operation; Dynamic (·) represents the gating decision function, and its specific calculation formula is shown in the following formula: (6); In the formula, F input Represents the input feature map; W 1. W 2 represents a learnable parameter; μ ( F input ) indicates to F input Perform channel average activation; σ ( F input )express F input Standard deviation; e ( F input )express F input Shannon entropy; b 1 and b 2 represents the threshold for controlling the activation of the hidden layer; S222, will F CBAM1 The data is fed into the encoder, which contains cubic 2D convolutions; after passing through the cubic 2D convolutions contained in the encoder, F CBAM1 The size is reduced to 0.5 times its original size, and the calculation formula is shown in equation (7): (7); In the above formula, Encoder Indicates encoder, F CBAM1 This represents the feature map after initial attention weighting. This represents the feature map after being re-encoded by the encoder. S223, after obtaining the encoder's output... Then, it is fed into the second DCBAM module, where it is attention-weighted again to obtain a second attention-weighted feature map. F CBAM2 The calculation formula is shown in equation (8): (8); In the above formula, Dynamic (·) represents the gating decision function; Channel attention2 This represents the weight of the attention for the second channel; Indicates to Perform channel attention weighting; Spatial attention2 Indicates the weights of the second spatial attention; Indicates to Perform spatial attention weighting; Indicates cumulative operation; This represents a feature-weighted operation; S224, will F CBAM2 The data is fed into the decoder. The decoder part contains a scale-invariant two-dimensional convolution and three upsampling operations. After the convolution in the decoder, the size of the feature becomes four times the original size. The calculation formula is shown in equation (9). (9); In the above formula, F CBAM2 This represents the feature map after a second attention weighting. Decoder Indicates decoder; This represents the feature map after decoding by the decoder; S225, after obtaining Then, after attention-weighted processing by the third DCBAM module at the output, the feature map that finally recovers the feature information of the small target is obtained. F CBAM3 The calculation formula is shown in equation (10): (10); In the formula, Dynamic (·) represents the gating decision function; Channel attention3 The weights representing the attention of the third channel; Indicates to Perform channel attention weighting; Spatial attention3 The weights represent the third spatial attention; Indicates to Perform spatial attention weighting; Indicates cumulative operation; This represents a feature-weighted operation; S226, In order to distinguish between the features extracted by the feature extraction module and the features in the real data samples, a discriminator network is also embedded at the end of the feature extraction module. The discriminator network includes a two-dimensional convolution and three downsampling operations, specifically: the features generated by the feature extraction module are processed by the discriminator network. F fake and real features F real The two vectors are fed into the discriminator network respectively. After obtaining the output vectors, the discriminant loss between them is calculated. The calculation formula is shown in Equation (11): (11); In the above formula, F fake This represents the features generated by the feature extraction module; F real || represents true features; F real and F fake Perform an OR operation between them; JUDG This represents the discriminator network; D output This represents the discrimination loss of the discriminator network's output. S23, Construct a geological feature recognition module, which combines the spectral information of the image and performs spatial attention, channel attention and self-attention on the spectral features of the image, and maps the multi-attention fused spectral features into image features; S24, the geological target feature construction module, geological feature recognition module, feature fusion module ASFF and FastestDet detection head are embedded into the backbone network CSPDarknet module to obtain the initial deep learning network model; S3, input the initial training samples into the initial deep learning network model for training, and obtain the trained deep learning network model; S4 inputs the image to be detected into the trained deep learning network model, which outputs the predicted location and type of the small geological target.

2. The method for visual detection of small targets in geological exploration scenarios according to claim 1, characterized in that: In step S22, the geological target feature construction module is constructed based on the feature extraction module built by the generative network and the two-stage kernel convolutional unit. The two-stage kernel convolutional unit consists of a spatial kernel convolution and a point kernel convolution, specifically represented by the following formula: (1); In the formula, F Features representing the input two-stage kernel convolutional units; PC This represents a two-order kernel convolutional unit; G Represents spatial kernel convolution; Z Represents a point kernel convolution; * indicates element-wise multiplication; In addition, the number of parameters of a two-stage kernel convolutional unit P DS and computational complexity C DS Represented as equations (2) and (3): P DS =D W ×D H ×M+ 1 × 1 ×N×M (2); C DS =(D W ×D H ×M+ 1 × 1 ×N×M)×F W ×F H (3); in, F W and F H These represent the length and width of the feature map extracted by the feature extraction module, respectively. M The number of feature map channels extracted by the feature extraction module; N This represents the number of convolution kernels; the size of the convolution kernel is... D W × D H × M ,in D W and D H These represent the length and width of the convolution kernel, respectively. The formula (4) for calculating the adversarial loss function of the generative adversarial network is: L = E x~Pr [ f w ( x )]- E x~Pg [ f w ( x )](4); in, f w ( x ) represents hyperparameters w The mean squared error function; E x~Pr Indicates from sampling point x~P r Obtain a sample x Expected value; E x~Pg Indicates from sampling point x~Pg Obtain a sample x Expected value L This represents the calculated adversarial loss value.

3. The method for visual detection of small targets in geological exploration scenarios according to claim 1, characterized in that: In steps S222 and S224, the two-dimensional convolution operation is augmented with BatchNorm2d normalization and LeakyReLU activation; in step S226, the two-dimensional convolution operation is augmented with InstanceNorm2d normalization and LeakyReLU activation.

4. The method for visual detection of small targets in geological exploration scenarios according to any one of claims 1 to 3, characterized in that: The geological target feature construction module also includes an output unit, which combines the outputs of the feature extraction module and the two-stage kernel convolution unit with the output of the feature fusion module ASFF and performs two-dimensional convolution operations.

5. The method for visual detection of small targets in a geological exploration scenario according to claim 1, characterized in that: In step S23, the construction process of the geological feature identification module is as follows: S231, firstly, the spectral data of the feature map is calculated, and spatial attention is used to capture the spatial dependencies between different locations in the spectrum, simultaneously focusing on the characteristics of different regions across the entire image space; specifically, the input feature map... F∈ R H×W×C Different spectral decomposition spatial attention methods are applied to focus on small target regions from different dimensions and perceive the characteristics of these regions. R H×W×C Representing feature maps, H and W These represent the length and width of the matrix, respectively. C The number of channels representing the matrix is ​​calculated using the formula shown in equation (12): F spatial_n = SpAttention ( G sw ( F ))(12); in, G sw Representative of feature maps F Calculate its spectral data to achieve the mapping from image features to spectral features; F spatial_n Represents the spatial attention spectral feature map, where the subscripts are... n This represents the dimension of spatial attention; SpAttention Represents the spectral feature map of the input. G sw ( F Spatial attention is performed in the corresponding different dimensions; S232, for feature maps F∈R H×W×C After calculating its spectral data, channel attention is applied to each channel of the spectral feature matrix to obtain the channel attention spectral feature map. F ch The calculation formula is shown in equation (13): F ch =Ch_Attention ( G sw ( F ))(13); in, Ch_Attention Indicates channel attention; G sw ( F () represents a spectral feature map; F ch This represents the channel attention spectral feature map; S233, for feature maps F∈R H×W×C After calculating its spectral data, self-attention is applied to all channels of the spectral feature matrix to obtain the self-attention spectral feature map. F self The calculation formula is shown in equation (14): F self =Self_Attention ( G sw ( F ))(14); in, Self Attention Indicates self-attention; G sw ( F () represents a spectral feature map; F self Represents the self-attention spectral feature map; S234, performs spectral adaptive fusion of the spatial attention spectral feature map, the channel attention spectral feature map, and the self-attention spectral feature map, and the calculation formula is shown in the following formula: (15); In the formula, F O This indicates the final feature obtained; F k ∈{ F spatial_n , F ch , F self }; MAP ( F k ) indicates that the spectral feature map F k Mapping spectral features to image features; W k , σ k , μ k This represents the weights, variances, and mean values ​​of the corresponding feature maps. ∑ indicates Euclidean operations; ∑ indicates accumulation operations.

6. The method for visual detection of small targets in a geological exploration scenario according to claim 4, characterized in that: In step S24, the geological target feature construction module, geological feature recognition module, and feature fusion module ASFF are embedded into the backbone network CSPDarknet module and combined with a FastestDet detection head to obtain the initial deep learning network model; the backbone network CSPDarknet includes six layers arranged sequentially; the construction process of the initial deep learning network model is as follows: S241, the output feature map of the fourth layer of the backbone network CSPDarknet is used as the input of the geological feature recognition module, and then the output of the geological feature recognition module and the output of the fourth layer are used as the input of the fifth layer of the backbone network CSPDarknet. S242, the outputs of the fourth, fifth, and sixth layers of the backbone network CSPDarknet are used as input features of the feature fusion module ASFF. Then, the feature fusion module ASFF performs fusion processing on the input features and outputs feature maps. The feature fusion module ASFF outputs three layers of feature maps with gradually decreasing scale: fused map one, fused map two, and fused map three. S243, the input to the geological target feature construction module is divided into two parts. The first part of the input is the output feature map of the fourth layer of the backbone network CSPDarknet; the second part of the input is the fusion map of the feature fusion module ASFF. The first part of the input serves as the input to the feature extraction module and the double-kernel convolutional unit, while the second part of the input, along with the outputs of the feature extraction module and the double-kernel convolutional unit, serves as the input to the output unit. The output of the output unit serves as the output of the geological target feature construction module. S244, the output of the feature fusion module ASFF and the geological target feature construction module, serves as the input of the FastestDet detection head; the FastestDet detection head outputs the category and location information of the target in the image.

7. The method for visual detection of small targets in a geological exploration scenario according to claim 6, characterized in that: The backbone network CSPDarknet has the following layers: the first layer is the Focus layer; the second layer is a 2D convolutional layer, normalization layer, and activation function layer; the third layer is a 2D convolutional layer, normalization layer, activation function layer, and convolutional layer layer; the fourth layer is a 2D convolutional layer, normalization layer, activation function layer, and convolutional layer layer; the fifth layer is a 2D convolutional layer, normalization layer, activation function layer, and convolutional layer layer layer; and the sixth layer is a 2D convolutional layer, normalization layer, activation function layer, spatial pyramid pooling layer, and convolutional layer layer.

Citation Information

Patent Citations

  • Underwater target identification method based on improved YOLOv8 algorithm

    CN120356084A

  • Fan blade defect detection method and system based on target detection model

    CN120431074A