Remote sensing open-pit mine detection method based on dual-branch structure and feature fusion mechanism

By constructing a remote sensing open-pit mine detection method based on a dual-branch structure and feature fusion mechanism, the detection accuracy problem under the influence of low-resolution images and complex terrain in remote sensing technology is solved, and efficient and accurate mine detection is achieved.

CN120279437BActive Publication Date: 2025-08-15CHONGQING INST OF GEOLOGY & MINERAL RESOURCES
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510757188.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-08-15
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

In the detection of open-pit mines, existing remote sensing technology has problems such as difficulty in clear extraction of mining area boundaries and land features caused by low-resolution images, uneven color distribution caused by changes in lighting conditions, and complex topography affecting detection accuracy. Traditional methods are inefficient and have high computational complexity.

Method used

The remote sensing open-pit mine detection method based on the dual-branch structure and feature fusion mechanism is adopted. By acquiring multi-source high-resolution remote sensing image data for preprocessing, a lightweight MobileNet V2 backbone network and feature extraction branches of dense generation module are built, and feature fusion and detection are performed in combination with cross attention and feature adaptive fusion module.

Benefits of technology

It improves the recognition accuracy of small target details and mining area boundaries in low-resolution images, reduces the computational complexity, enhances the adaptability and detection accuracy of the model, and is suitable for multi-scale diversified remote sensing scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279437B_ABST
    Figure CN120279437B_ABST
Patent Text Reader

Abstract

The present invention provides a remote sensing open-pit mining area detection method based on a dual-branch structure and a feature fusion mechanism, comprising: obtaining multi-source high-resolution remote sensing image data of an open-pit mining area, and performing preprocessing such as atmospheric correction, radiation calibration, color correction, image registration, and cropping operations on the data to obtain time-series remote sensing image data; using a deep learning algorithm to construct an open-pit mining area detection model based on the dual-branch structure and feature fusion mechanism, and training the model using the time-series remote sensing image data to obtain a remote sensing image detection model, wherein the dual-branch structure includes a feature extraction branch and a feature generation branch, and the feature fusion mechanism includes a cross-attention fusion module and a feature adaptive fusion module; and inputting the remote sensing image data to be tested into the remote sensing image detection model to obtain detection results of the open-pit mining area. The present invention improves the recognition accuracy of small target details and mining area boundaries in low-resolution images, while balancing computational efficiency and detection accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image processing, and in particular to a remote sensing open-pit mine area detection method based on a dual-branch structure and a feature fusion mechanism. Background Art

[0002] The traditional model of mineral resource supervision relies primarily on manual inspections and on-site law enforcement. This approach is not only time-consuming and labor-intensive, but is also significantly affected by factors such as terrain and climate, making it difficult to achieve efficient and accurate monitoring of large areas. With the rapid development of remote sensing technology and the continuous improvement of aerospace detection capabilities, the ability to acquire high-resolution remote sensing imagery has been significantly enhanced, greatly promoting the widespread application of remote sensing data in multiple fields, such as land resource management, urban planning, crop management, environmental protection, and natural disaster monitoring. It has also brought new opportunities for the dynamic monitoring of mineral resources. Remote sensing images have the characteristics of large-scale coverage and periodic updates. They can continuously obtain surface information in mining areas, providing accurate and efficient data support for the dynamic monitoring of mining activities. Currently, the use of high-resolution remote sensing imagery for automatic detection and identification of open-pit mines has become an important technical means for mineral resource management.

[0003] However, due to the unique imaging conditions of remote sensing images and the complex geographical environment of open-pit mines, current remote sensing detection of open-pit mines still faces numerous challenges. First, because the resolution of remote sensing images is limited by factors such as the imaging equipment and shooting height, many images still suffer from low resolution. This results in blurred surface texture information in the mining area, making it difficult to clearly extract mining area boundaries and ground features. Second, the variable lighting conditions and imaging angles of remote sensing images can easily lead to uneven color distribution, resulting in low contrast between the mining area and the surrounding environment, which increases the difficulty of target recognition. Furthermore, open-pit mines typically have complex topographical features, and the ground features of different mineral types vary significantly, making their boundaries, shapes, and spatial distribution difficult to accurately model and describe. These factors severely affect the detection accuracy of existing methods. Therefore, how to effectively overcome the problems of low remote sensing image resolution, blurred texture information, and uneven colors, accurately analyze the structural characteristics of the mining area (such as boundaries, shapes, and spatial distribution), and improve the accuracy of target recognition are important issues that need to be addressed in current remote sensing open-pit mine detection research.

[0004] In the early days, the identification of open-pit mines relied primarily on manual visual interpretation and annotation. However, this approach is not only labor-intensive and time-consuming, but also highly reliant on expert experience, subjectivity, inefficiency, and unstable accuracy. Furthermore, factors such as remote sensing image quality, weather conditions, and lighting variations often make it difficult for manual interpretation to accurately define mine boundaries, resulting in significant errors in detection results. With the advancement of machine learning technology, researchers have gradually explored the application of machine learning methods to remote sensing mine detection and continuously optimized algorithms to improve detection accuracy. In machine learning-based methods, mine detection is typically considered a pixel-level classification task in remote sensing imagery. This involves classifying each pixel as either a mine or non-mine area, and automatically identifying the area using machine learning classification algorithms. While machine learning methods improve detection efficiency and reduce subjective errors to a certain extent compared to traditional manual interpretation methods, early machine learning approaches still face numerous challenges. For example, efficiently processing and extracting key features from massive amounts of remote sensing data remains a challenge. In addition, the topography of the mining area is complex, and the spectral characteristics of different bands vary greatly. A single machine learning model is difficult to accurately depict the boundaries and spatial distribution relationships of the mining area, resulting in limited detection accuracy.

[0005] In recent years, with the rapid development of deep learning, deep learning-based methods have achieved remarkable success in image processing. Leveraging their powerful nonlinear mapping capabilities and end-to-end learning mechanisms, they have significantly improved the performance of object detection, segmentation, and recognition tasks. Compared to traditional methods that rely on manually designed features, deep learning-based methods can automatically learn effective features through large-scale data training, eliminating the subjectivity of feature selection while improving the model's adaptability and generalization capabilities. By constructing multi-layer neural networks, deep learning models can gradually extract low-level to high-level semantic information, thereby more accurately understanding the structure of objects in an image and their relationship to the background. Furthermore, the introduction of attention mechanisms further enhances the model's ability to model global information, enabling it to perform better in complex environments and long-range dependencies. However, due to the complex surface environment in mining areas, object boundaries often exhibit gradual transitions, making it difficult for detection models to accurately distinguish between mining areas and surrounding objects, leading to significant false detections and missed detections. Furthermore, due to the limitations of remote sensing image acquisition, much optical remote sensing data has low resolution, and each pixel may contain multiple object types, making it even more difficult to effectively extract texture and structural information. Furthermore, lighting variations, shadow interference, and the uneven color distribution of different landforms within mining areas limit the generalization capabilities of traditional deep learning methods across diverse scenarios. Currently, most deep learning methods rely on deepening network structures or increasing trainable parameters to improve detection performance. However, this often leads to significant increases in computational complexity and storage requirements, limiting the models' practical application in resource-constrained environments.

[0006] Therefore, for the remote sensing detection task of open-pit mines, there is an urgent need for a lightweight remote sensing open-pit mine detection method that can improve the recognition accuracy of small target details and mine boundaries in low-resolution images, while taking into account both computational efficiency and detection accuracy. Summary of the Invention

[0007] Based on this, it is necessary to provide a remote sensing open-pit mine area detection method based on a dual-branch structure and feature fusion mechanism to address the above technical problems.

[0008] A remote sensing open-pit mine detection method based on a dual-branch structure and a feature fusion mechanism comprises the following steps: acquiring multi-source high-resolution remote sensing image data of an open-pit mining area; preprocessing the remote sensing image data to obtain time-series remote sensing image data, wherein the preprocessing includes atmospheric correction, radiometric calibration, color correction, image registration, and cropping operations; constructing an open-pit mine detection model based on the dual-branch structure and feature fusion mechanism using a deep learning algorithm, and training the model using the time-series remote sensing image data to obtain a remote sensing image detection model, wherein the dual-branch structure includes a feature extraction branch and a feature generation branch, and the feature fusion mechanism includes a cross-attention fusion module and a feature adaptive fusion module; wherein the feature extraction branch uses a lightweight MobileNet V2 as a backbone network, and removes the final pooling layer and fully connected layer to extract multi-scale preliminary features from the remote sensing image data; the feature generation branch is composed of multiple dense generation modules to generate multi-scale fine-grained features from the remote sensing image data to enhance the learning of mining area edges and texture details; The cross-attention fusion module is used to supplement the multi-scale fine-grained features into the multi-scale preliminary features; the feature adaptive fusion module adopts a ladder structure for stacking, and performs the fusion of adjacent scale features and the interaction of non-adjacent scale features during the decoding process to fully capture the detailed information of the mining area target; the remote sensing image data to be tested is input into the remote sensing image detection model to obtain the detection results of the open-pit mining area.

[0009] In one embodiment, the MobileNet V2 network includes an input layer, an initial convolution layer, multiple inverted residual structure modules, a global average pooling layer and a fully connected layer; wherein, the input layer is used to receive preprocessed time series remote sensing image data; the initial convolution layer uses a convolution kernel of size 3×3 and a step size of 2 to perform preliminary feature extraction on the input image; the inverted residual structure module includes an expansion layer, a depthwise separable convolution layer and a projection layer, wherein the expansion layer is a 1×1 convolution operation for expanding the channel dimension of the input feature; the depthwise separable convolution layer is composed of a 3×3 depthwise convolution and a projection layer. 1×1 point-by-point convolution, where depthwise convolution is used to perform spatial convolution operations on each channel independently, and point-by-point convolution is used to integrate information between channels; the projection layer uses 1×1 convolution to compress high-dimensional features back to the original low-dimensional space, and fuse them with the input features through jump connections; multiple inverted residual structure modules are stacked with different numbers of channels and resolutions to construct a deep feature extraction network; the global average pooling layer is used to compress the spatial dimension of the feature map; the fully connected layer is used to map the global features to the classification results or detection outputs.

[0010] In one embodiment, the dense generation module includes three residual dense blocks, each of which includes five 3×3 convolutional layers connected in sequence, and Leaky ReLU is used as the activation function after each convolution operation, and no batch normalization layer is included; the convolutional layers are jump-connected by splicing the channel dimension, and the residual dense blocks are feature-fused by element-by-element addition, and a residual scaling coefficient is introduced to scale the fusion result; the structural end of the dense generation module uses a 3×3 convolutional layer with a stride of 2 for downsampling to keep the size of the feature map consistent with the feature extraction branch; wherein, the formula of the dense generation module is:

[0011] ;

[0012] Where, Represents a convolution operation with a kernel size of 3 and a stride of 1; Indicates the first The output result of the convolution operation; Indicates that in the current residual dense block, the features of the previous layers are concatenated in the channel dimension before the i-th convolution; Represent the output results of the three residual dense blocks respectively; Represents the Leaky ReLU activation function.

[0013] In one embodiment, the cross-attention fusion module includes a spatial attention block and a channel attention block; through the spatial attention block and the channel attention block, the multi-scale preliminary features and the multi-scale fine-grained features are respectively subjected to attention operations in the spatial dimension and the channel dimension to obtain a channel attention sequence and a spatial attention feature map, and a cross-matrix multiplication operation is performed to obtain two weight matrices; based on the two weight matrices, the multi-scale preliminary features and the multi-scale fine-grained features are respectively subjected to a re-weighting operation of element-by-element multiplication; the re-weighted features are subjected to a feature fusion operation of element-by-element addition to obtain a fused feature; the fused feature is adjusted through the depthwise separable convolutional layer to obtain an output feature after fusing the two branch features.

[0014] In one embodiment, the channel attention block is used to perform global maximum pooling and global average pooling on the input feature map in the spatial dimension to obtain two different sets of channel information sequences; the two sets of channel information sequences are input into the multi-layer perceptron through the channel attention block for feature reconstruction to extract high-level channel features; the two sets of channel features after feature reconstruction are fused through element-level addition to obtain an output channel attention sequence containing rich channel information of the input feature map; wherein the formula of the channel attention block is:

[0015] ;

[0016] Where, and Represent global average pooling and global maximum pooling operations respectively, MLP represents multi-layer perceptron, represents the input features, Represents the output sequence of the channel attention module.

[0017] In one embodiment, the spatial attention block is used to enhance the semantic representation capability of the feature map; the spatial attention block performs global maximum pooling and global average pooling operations on the channel dimension of the input feature map, and combines 1×1 convolution to compress the channel dimension of the input feature map to generate three spatial information matrices; the three spatial information matrices are spliced on the channel dimension and compressed through convolution operation to obtain the output spatial information matrix , used to guide the feature enhancement of important spatial regions, the formula is:

[0018] ;

[0019] ;

[0020] ;

[0021] ;

[0022] Where, and Represent the maximum pooling operation and average pooling operation in the channel dimension respectively; Represents the output spatial information matrix of the channel attention module; Represents a point convolution operation with a convolution kernel size of 1; Indicates that the concatenation operation is performed on the channel dimension.

[0023] In one embodiment, the feature adaptive fusion module is used to receive two sets of feature maps at adjacent scales, namely low-level feature maps With high-level feature maps ; The high-level feature map is subjected to bilinear interpolation upsampling processing through the feature adaptation module to make its spatial resolution consistent with that of the low-level feature map; the upsampled high-level feature map is subjected to 3×3 convolution to make it consistent with the number of channels of the low-level feature map and the low-level feature map. Perform element-level addition operation to perform preliminary feature fusion. The formula is:

[0024] ;

[0025] Where, Represents bilinear interpolation upsampling operation; for fusion features 3×3 dilated convolution operations with dilation rates of 1, 3, and 5 are used respectively to obtain feature representations under three different receptive fields, which are then concatenated in the channel dimension. 3×3 convolution is used to restore the channel dimension. The process is as follows:

[0026] ;

[0027] ;

[0028] Where, The expansion rate is Convolution operation; for low-level feature maps , feature compression is performed through a 1×1 convolution combined with a Sigmoid activation function to obtain a pixel-level guidance mask , whose expression is:

[0029] ;

[0030] By setting all elements of the matrix to 1 minus Get the reverse mask ,in, and are used to represent the context information of the target area and the background area respectively; and After splicing in the channel dimension, it is input into the 1×1 convolution and Sigmoid activation function to generate a set of Pixel weight matrix with the same resolution and number of channels , used for The features are recalibrated to obtain the self-guided correction features , the calculation formula is:

[0031] ;

[0032] Where, Indicates the mask Take the inverse, that is, from the matrix Subtract The output of the feature adaptive fusion module is composed of the aligned and corrected high-level features and the low-level features after self-guided correction After weighted fusion and feature integration through 3×3 convolution, the formula is:

[0033] ;

[0034] Where, is the output of the feature adaptive fusion module.

[0035] In one embodiment, after the open-pit mine area detection model based on the dual-branch structure and feature fusion mechanism is constructed by using the deep learning algorithm and trained with the time series remote sensing image data to obtain the remote sensing image detection model, before the remote sensing image detection model is input with the remote sensing image data to obtain the detection result of the open-pit mine mining area, the method further includes: converting the binary cross entropy loss With Dice loss Perform weighted fusion and construct a joint loss function. The formula is:

[0036] ;

[0037] ;

[0038] ;

[0039] Where, and Respectively and The maximum value of and Represents the position in the image The predicted value and the true label at Represents the cross entropy loss function, symbol express norm; the feature adaptive fusion module is stacked in a step-by-step manner to generate three output results that are consistent with the resolution of the lowest layer feature map; the loss values of the three output results are calculated by the joint loss function, and the three calculated loss values are assigned different weight coefficients and then weighted summed to obtain the final loss value. The formula is:

[0040] ;

[0041] ;

[0042] Where, 、 and They represent the loss values calculated from the output results of the three lowest-level feature adaptive fusion modules, 、 and are the weight coefficients assigned to the corresponding loss items, is the final loss value.

[0043] Compared with the existing technology, the advantages and beneficial effects of the present invention are as follows: by acquiring multi-source high-resolution remote sensing image data covering the entire range of the open-pit mining area from different satellites, and performing pre-processing such as atmospheric correction, radiation calibration, color correction, image registration and cropping operations on it, time-series remote sensing image data is obtained, and the contrast of the image is improved, making it more conducive to model detection; a deep learning algorithm is used to construct an open-pit mining area detection model based on a dual-branch structure and a feature fusion mechanism, and the remote sensing image detection model is obtained by training with time-series remote sensing image data. The dual-branch structure includes a feature extraction branch and a feature generation analysis, and the feature fusion mechanism includes a cross-attention fusion module and a feature adaptive fusion module. The trained remote sensing image detection model is used to detect the remote sensing image data to be tested, and the detection results of the open-pit mining area are obtained; combined with the dual-branch structure and feature fusion mechanism, it can deeply mine the structural information and detail features in the remote sensing image, and strengthen the feature fusion expression, thereby improving the model's ability to recognize complex landforms, ensuring the recognition accuracy of small target details and mining area boundaries in low-resolution images, thereby improving the accuracy and robustness of remote sensing image detection, and can be applied to target detection tasks in multi-scale and diversified remote sensing scenarios. In addition, the present invention has a low number of parameters and can achieve efficient calculation, thereby ensuring both detection accuracy and calculation efficiency, and improving the deployment efficiency and practicality of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 1 is a flow chart of a remote sensing open-pit mine area detection method based on a dual-branch structure and a feature fusion mechanism in one embodiment;

[0045] Figure 2 Schematic diagram of the architecture of a remote sensing image detection model in one embodiment;

[0046] Figure 3 Schematic diagram of the structure of a cross-attention fusion module in one embodiment;

[0047] Figure 4 Schematic diagram of the structure of a feature adaptive fusion module in one embodiment;

[0048] Figure 5 is a schematic structural diagram of a dense generation module in one embodiment;

[0049] Figure 6 The figure is a schematic diagram of a visual comparison between the present invention and other existing methods on a data set in one embodiment. DETAILED DESCRIPTION

[0050] Before describing the specific embodiments of the present invention, the overall concept of the present invention is described as follows:

[0051] The present invention is mainly developed based on the remote sensing data processing process. Traditional remote sensing data processing methods have poor recognition performance for small target details and mining area boundaries when facing low-resolution image detection, low detection accuracy, or complex model structure, which is not conducive to practical application.

[0052] Therefore, the present invention proposes a remote sensing open-pit mine area detection method based on a dual-branch structure and a feature fusion mechanism, obtains multi-source high-resolution remote sensing image data of the open-pit mine sampling area, forms a remote sensing image data set, and performs preprocessing such as atmospheric correction, radiation calibration, color correction, image registration and cropping operations on the remote sensing image data to improve image quality and subsequent modeling effects; a deep learning algorithm is used to construct an open-pit mine area detection model based on a dual-branch structure and a feature fusion mechanism, and the remote sensing image detection model is obtained by training with time-series remote sensing image data. The dual branches include a feature extraction branch and a feature generation branch, and the feature fusion mechanism of the cross-attention fusion module and the feature adaptive fusion module are introduced to achieve efficient feature interaction and fusion between the dual branches, thereby enhancing The ability to extract target features of mining areas in complex remote sensing images is achieved by using the trained remote sensing image detection model to detect and process the remote sensing image data to be tested, and accurate identification results of open-pit mining areas are obtained. Combined with the dual-branch structure and feature fusion mechanism, it can deeply mine the structural information and detail features in remote sensing images, and strengthen the feature fusion expression, thereby improving the model's ability to identify complex landforms, ensuring the recognition accuracy of small target details and mining area boundaries in low-resolution images, thereby improving the accuracy and robustness of remote sensing image detection, and can be applied to target detection tasks in multi-scale and diversified remote sensing scenarios. In addition, the present invention has a low number of parameters and can achieve efficient calculation, thereby ensuring both detection accuracy and calculation efficiency, and improving the deployment efficiency and practicality of the model.

[0053] After introducing the overall concept of the present invention, in order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below through specific embodiments in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0054] In one embodiment, Figure 1 As shown, a remote sensing open-pit mine area detection method based on a dual-branch structure and a feature fusion mechanism is provided, comprising the following steps:

[0055] Step S110: Acquire multi-source high-resolution remote sensing image data of the open-pit mine mining area.

[0056] Specifically, multi-source, high-resolution remote sensing imagery covering the open-pit mining area is acquired through various types of remote sensing satellites. This remote sensing imagery data, collected by multiple satellite platforms, boasts high spatial resolution and imaging frequency, adapting to imaging requirements in diverse terrain conditions. The acquired remote sensing imagery covers a wide range of open-pit mining areas (such as sandstone, limestone, and shale), ensuring the detection model's versatility and adaptability for identifying and extracting diverse mineral types.

[0057] Step S120 , preprocessing the remote sensing image data to obtain time series remote sensing image data, wherein the preprocessing includes atmospheric correction, radiometric calibration, color correction, image registration, and cropping operations.

[0058] Specifically, the acquired satellite remote sensing image data undergoes systematic preprocessing. First, atmospheric correction is performed to remove the effects of atmospheric scattering and absorption. Next, radiometric calibration is performed to unify the physical response values of imaging data from different sensors. Color correction is then performed to improve the color consistency and visual quality of the images. Precise geometric registration is performed on images from different satellite platforms or imaging times to ensure spatial position consistency. Finally, image cropping is performed based on the boundaries of the open-pit mine area to extract image segments of the target area. Through these preprocessing operations, time-series remote sensing image data is generated for subsequent model training, significantly improving the quality and consistency of remote sensing data and laying the foundation for accurate detection and feature extraction in subsequent models.

[0059] In step S130, a deep learning algorithm is used to construct an open-pit mine detection model based on a dual-branch structure and a feature fusion mechanism, and the model is trained using time-series remote sensing image data to obtain a remote sensing image detection model. The dual-branch structure includes a feature extraction branch and a feature generation branch, and the feature fusion mechanism includes a cross-attention fusion module and a feature adaptive fusion module.

[0060] Among them, the feature extraction branch uses a lightweight MobileNet V2 as the backbone network, and removes the final pooling layer and fully connected layer to extract multi-scale preliminary features from the remote sensing image data; the feature generation branch is composed of multiple dense generation modules to generate multi-scale fine-grained features from the remote sensing image data, enhancing the learning of mining area edges and texture details; the cross-attention fusion module is used to supplement the multi-scale fine-grained features to the multi-scale preliminary features; the feature adaptive fusion module is stacked with a ladder structure, and performs the fusion of adjacent scale features and the interaction of non-adjacent scale features during the decoding process to comprehensively capture the detailed information of the mining area targets.

[0061] Specifically, a deep learning algorithm is used to construct an open-pit mine detection model based on a dual-branch structure and feature fusion mechanism, and pre-processed time series remote sensing image data is used for training. After training, a remote sensing image detection model is obtained, such as Figure 2 As shown in the figure, it is used for open-pit mine inspection and is suitable for accurate open-pit mine inspection in complex environments. The dual-branch structure includes a feature extraction branch and a feature generation branch, and the feature fusion mechanism includes a Cross-Attention Fusion Module (CAFM) and a Feature Adaptive Fusion Module (FAFM).

[0062] like Figure 2 As shown in the figure, it is a structural diagram of the remote sensing image detection model. The input image is processed by the feature extraction branch and the feature generation branch respectively. The feature extraction branch is used to extract features through stage1-stage5 and input into the cross attention fusion module for feature fusion, so that the model can fully obtain local detail features while ensuring the global semantic understanding ability; at the same time, the dense generation module (DGM) of the feature generation branch is used to perform efficient interaction between features at different levels of features. The fused features are input into the step-by-step feature adaptation module to fuse adjacent scale features and interact with non-adjacent scale features. Finally, the detection results are output, and the output detection results can be fed back and supervised by the ground truth to ensure the accuracy of the model, so that the model can fully capture the detailed information of the mining area target. The remote sensing image detection model obtained based on the above structure has high detection accuracy and strong generalization ability.

[0063] The system employs a dual-branch structure consisting of a feature extraction branch and a feature generation branch. The feature extraction branch extracts preliminary multi-scale features from acquired multi-source remote sensing imagery data. The branch uses the lightweight convolutional neural network MobileNet V2 as its backbone, removing the final pooling and fully connected layers. This reduces the computational complexity of the model while maintaining high feature extraction capabilities, improving the efficiency and deployment flexibility of the detection system. MobileNet V2 boasts a compact structure, a small number of parameters, and fast computational speed, making it particularly suitable for deployment on devices with limited computing resources.

[0064] The feature generation branch, composed of multiple dense generation modules, is used to perform fine-grained texture restoration and edge information enhancement on the initial image features, obtaining multi-scale fine-grained features. This allows for fine-grained texture restoration and edge information enhancement of multi-scale preliminary features, improving the detection capabilities of small targets and boundary structures. The feature generation branch can reconstruct fine-grained detail information from low-quality remote sensing imagery, compensating for information loss caused by low resolution during feature extraction. It can also enhance the texture features of the target area, making the detection results more accurate. Multiple dense generation modules enhance feature transmission and reuse capabilities through a large number of jump connections, thereby strengthening the ability to generate detailed features and mitigating the impact of low resolution remote sensing imagery.

[0065] The dual-branch structure can fully utilize the advantages of different branches, improve the classification and recognition capabilities of the model, enhance the robustness of the model, and enable it to better handle various input data. The multi-branch structure can extract multiple features from the same input, increase feature diversity, alleviate the overfitting problem of the model, and adapt to different task requirements.

[0066] In order to effectively supplement the fine-grained features extracted by the feature generation branch to the preliminary features obtained by the feature extraction branch, and improve the integrity and discrimination ability of the feature representation, a cross-attention fusion module is designed. Through the synergistic effect of spatial attention and channel attention, the features of the two branches are weighted and cross-fused to enhance the feature expression ability, strengthen the correlation and complementarity between the two types of features, and achieve more accurate feature fusion. In order to make full use of the acquired multi-scale features and make full use of multi-scale information in the decoding process, the present invention also designs a feature adaptive fusion module to gradually fuse high-level features to low-level features in the decoding process to improve the accuracy of detection. Therefore, the feature fusion mechanism of the present invention includes a cross-attention fusion module and a feature adaptive fusion module.

[0067] The structure of the cross attention fusion module is as follows Figure 3 As shown in the figure, it is used to supplement multi-scale fine-grained feature information with multi-scale preliminary features, enhancing the semantic expression ability of features. The cross-attention fusion mechanism can dynamically focus on the correlation between different modalities. Through the cross-attention mechanism, key features are highlighted and redundant information is printed. The semantic information of the feature extraction branch is effectively integrated with the fine-grained features of the feature generation branch, thereby enhancing feature representation ability, achieving efficient utilization of dual-branch features, and improving detection accuracy.

[0068] The structure of the feature adaptation module is as follows Figure 4 As shown, during the decoding process, dilated convolution and adaptive guidance mechanisms are used to aggregate high-level features into low-level features, achieving a complementary fusion of detail information and semantic information, fully utilizing multi-level features, thereby improving detection performance and generating feature maps with rich semantic expression capabilities. In the remote sensing image detection model, a ladder structure is used to stack feature adaptation modules layer by layer, achieving the fusion of adjacent scale features and the interaction of non-adjacent scale features to fully capture the detailed information of mining targets, ensuring that the final feature map can retain complete and accurate semantic information, thereby improving detection accuracy.

[0069] Through feature adaptive fusion and cross-attention fusion, the features of different modalities are weighted and fused. The weights are dynamically determined by the cross-attention mechanism to reflect the importance of different features. Combining the cross-attention fusion mechanism and the feature adaptive fusion mechanism can enhance the feature expression ability, improve the accuracy and robustness of the model, and can be applied to a variety of multimodal tasks.

[0070] In one embodiment, the MobileNet V2 network includes an input layer, an initial convolution layer, multiple inverted residual structure modules, a global average pooling layer, and a fully connected layer; wherein the input layer is used to receive preprocessed time series remote sensing image data; the initial convolution layer uses a convolution kernel of size 3×3 and a step size of 2 to perform preliminary feature extraction on the input image; the inverted residual structure module includes an expansion layer, a depthwise separable convolution layer, and a projection layer, wherein the expansion layer is a 1×1 convolution operation for expanding the channel dimension of the input feature; the depthwise separable convolution layer consists of a 3×3 depthwise convolution and The network is composed of 1×1 point-by-point convolutions, where depthwise convolution is used to perform spatial convolution operations on each channel independently, and point-by-point convolution is used to integrate information between channels. The projection layer uses 1×1 convolution to compress high-dimensional features back to the original low-dimensional space and fuse them with the input features through jump connections. Multiple inverted residual structure modules are stacked with different numbers of channels and resolutions to construct a deep feature extraction network. The global average pooling layer is used to compress the spatial dimension of the feature map. The fully connected layer is used to map the global features to the classification results or detection outputs.

[0071] Specifically, the MobileNet V2 network consists of an input layer, an initial convolutional layer, multiple inverted residual architecture modules, a global average pooling layer, and fully connected layers. The input layer receives preprocessed time-series remote sensing image data. The initial convolutional layer uses a 3×3 kernel with a stride of 2 to perform preliminary feature extraction on the input image and reduce the spatial resolution, thereby improving the computational efficiency of the subsequent deep network.

[0072] The inverted residual structure module is the core structure of MobileNet V2, mainly including the expansion layer, the depthwise separable convolution layer, and the projection layer. Among them, the expansion layer is a 1×1 convolution operation, which is used to expand the channel dimension of the input features to improve the feature expression capability; the depthwise separable convolution layer consists of a 3×3 depthwise convolution and a 1×1 pointwise convolution. The former performs spatial convolution operations on each channel independently, and the latter is used to integrate information between channels. This design significantly reduces the computational complexity of the model and improves the efficiency of spatial feature extraction; the projection layer also uses 1×1 convolution to compress high-dimensional features back to the original low-dimensional space, and fuses them with the input features through jump connections to form a complete inverted residual connection structure. Multiple inverted residual structure modules are stacked with different numbers of channels and resolutions to construct a deep feature extraction network.

[0073] In addition, the feature extraction branch also includes a global average pooling layer, which is used to compress the spatial dimension of the high-dimensional feature map and extract global semantic information; finally, the global features are mapped to the task output space through the fully connected layer, and the final detection results are output.

[0074] By introducing MobileNet V2 as the backbone network of the feature extraction branch, we can effectively control the model size and computing resource consumption while maintaining high detection accuracy, and have good practicality and deployment adaptability.

[0075] In one embodiment, the dense generation module includes three residual dense blocks, each of which includes five 3×3 convolutional layers connected once, with Leaky ReLU used as the activation function after each convolution operation, and does not include a batch normalization layer; the convolutional layers are jump-connected by splicing the channel dimension, and the residual dense blocks are feature-fused by element-by-element addition, and a residual scaling factor is introduced to scale the fusion result; the structural end of the dense generation module uses a 3×3 convolutional layer with a stride of 2 for downsampling to keep the size of the feature map consistent with the feature extraction branch; wherein, the formula of the dense generation module is:

[0076] ;

[0077] Where, Represents a convolution operation with a kernel size of 3 and a stride of 1; Indicates the first The output result of the convolution operation; Indicates that in the current residual dense block, the features of the previous layers are concatenated in the channel dimension before the i-th convolution; Represent the output results of the three residual dense blocks respectively; Represents the Leaky ReLU activation function.

[0078] Specifically, if Figure 5 As shown in the figure, each dense generation module consists of three sequentially stacked residual dense blocks (RDBs). Each residual dense block contains five sequentially connected 3×3 convolutional layers. The convolutional layers are connected by skip connections through channel-wise splicing to fully realize cross-layer feature transfer and reuse of low-level information. The Leaky ReLU function is used as the activation function after each convolution operation to enhance nonlinear expression capabilities. At the same time, the entire structure does not introduce batch normalization layers to reduce redundant computation during the training phase and improve the model's adaptability in complex remote sensing scenarios.

[0079] Within the residual dense block, the input of each convolutional layer is the concatenation of the outputs of all previous layers, forming a typical densely connected structure. Feature fusion is achieved between the three residual dense blocks using element-by-element addition to improve overall expressiveness. To control the scale of the fused features, a residual scaling factor β is introduced after feature fusion to scale the fused result, effectively suppressing gradient explosion and enhancing the stability of the training process.

[0080] In addition, in order to achieve spatial scale alignment with the feature extraction branch, the dense generation module sets a convolutional layer with a stride of 2 and a kernel size of 3×3 at the end of the structure to downsample the feature map so that its size is consistent with the output of the feature extraction branch, facilitating subsequent multi-branch feature fusion operations.

[0081] Through the collaborative design of a large number of jump connections and dense residual structures, the dense generation module achieves the effective fusion of low-level texture information and high-level semantic information in remote sensing images, strengthens the ability to preserve edges and local structures, and thus improves the overall model's recognition ability and detection accuracy for complex pattern structures, enhancing the robustness of the model in different types of open-pit mining scenarios.

[0082] In one embodiment, the cross-attention fusion module includes a spatial attention block and a channel attention block; through the spatial attention block and the channel attention block, attention operations are performed on the multi-scale preliminary features and the multi-scale fine-grained features in the spatial dimension and the channel dimension respectively to obtain a channel attention sequence and a spatial attention feature map, and a cross-matrix multiplication operation is performed to obtain two weight matrices; based on the two weight matrices, a re-weighting operation of element-by-element multiplication is performed on the multi-scale preliminary features and the multi-scale fine-grained features respectively; a feature fusion operation of element-by-element addition is performed on the re-weighted features to obtain a fused feature; the fused feature is adjusted through a depth-wise separable convolutional layer to obtain an output feature after fusing the two branch features.

[0083] Specifically, if Figure 3 As shown in the figure, the cross-attention fusion module includes the spatial attention block (SAB) and the channel attention block (CAB). Through the synergy of spatial attention and channel attention, the correlation and complementarity between the two types of features are enhanced to achieve more accurate feature fusion.

[0084] The multi-scale preliminary features output by the feature extraction branch and the multi-scale fine-grained features output by the feature generation branch are subjected to attention operations in the spatial and channel dimensions, respectively. The resulting channel attention sequence and spatial attention feature map are then cross-matrix multiplied to obtain two weight matrices. Based on these two weight matrices, the multi-scale preliminary features and the multi-scale fine-grained features are element-wise multiplied to reweight them. The weighted features are then element-wise added to achieve feature fusion. Finally, the fused features are adjusted through a depthwise separable convolutional layer to obtain the final output feature that fuses the features of the two branches.

[0085] The channel attention block is used to perform global maximum pooling and global average pooling on the input feature map in the spatial dimension to obtain two different sets of channel information sequences. The two sets of channel information sequences are input into the multilayer perceptron (MLP) through the channel attention block for feature reconstruction and extraction of high-level channel features. The two sets of channel features after feature reconstruction are fused through element-wise addition to obtain an output channel attention sequence containing rich channel information of the input feature map. The formula of the channel attention block is:

[0086] ;

[0087] Where, and Represent global average pooling and global maximum pooling operations respectively, MLP represents multi-layer perceptron, represents the input features, Represents the output sequence of the channel attention module.

[0088] Specifically, if Figure 4 As shown in the figure, the channel attention block includes performing global max pooling (GAP) and global average pooling (GAP) on the input feature map in the spatial dimension to obtain two sets of different channel information sequences; the two sets of channel information sequences are input into the multi-layer perceptron for feature reconstruction to extract high-level channel features; the two sets of reconstructed channel features are fused by element-level addition to obtain an output channel attention sequence containing rich channel information of the input feature map.

[0089] Among them, the spatial attention block is used to enhance the semantic representation ability of the feature map. The spatial attention block performs global maximum pooling and global average pooling operations on the channel dimension of the input feature map respectively, and combines the global maximum pooling and global average pooling operations on the channel dimension of the input feature map respectively. The channel dimension of the input feature map is compressed by combining 1×1 convolution to generate three spatial information matrices; the three spatial information matrices are spliced in the channel dimension and compressed by convolution operation to obtain the output spatial information matrix , used to guide the feature enhancement of important spatial regions, the formula is:

[0090] ;

[0091] ;

[0092] ;

[0093] ;

[0094] Where, and Represent the maximum pooling operation and average pooling operation in the channel dimension respectively; Represents the output spatial information matrix of the channel attention module; Represents a point convolution operation with a convolution kernel size of 1; Indicates that the concatenation operation is performed on the channel dimension.

[0095] Specifically, the spatial attention block in the cross-attention fusion module is used to enhance the semantic representation ability of the feature map. Its processing flow includes: first, performing global maximum pooling and global average pooling operations on the channel dimension of the input feature map respectively, and combining 1×1 convolution to compress the channel dimension of the input feature map to generate three spatial information matrices; then, the three spatial information matrices are spliced on the channel dimension and compressed through another convolution operation to obtain the output spatial information matrix, which is used to guide the feature enhancement of important spatial areas, thereby improving the expressiveness of the feature fusion results.

[0096] In one embodiment, the feature adaptive fusion module is used to receive two sets of feature maps at adjacent scales, namely, low-level feature maps With high-level feature maps ; The high-level feature map is upsampled by bilinear interpolation through the feature adaptation module to make its spatial resolution consistent with the low-level feature map; the upsampled high-level feature map is convolved by 3×3 to make it consistent with the number of channels of the low-level feature map and the low-level feature map. Perform element-level addition operation to perform preliminary feature fusion. The formula is:

[0097] ;

[0098] Where, Represents bilinear interpolation upsampling operation; for fusion features 3×3 dilated convolution operations with dilation rates of 1, 3, and 5 are used respectively to obtain feature representations under three different receptive fields, which are then concatenated in the channel dimension. 3×3 convolution is used to restore the channel dimension. The process is as follows:

[0099] ;

[0100] ;

[0101] Where, The expansion rate is Convolution operation; for low-level feature maps , feature compression is performed through a 1×1 convolution combined with a Sigmoid activation function to obtain a pixel-level guidance mask , whose expression is:

[0102] ;

[0103] By setting all elements of the matrix to 1 minus Get the reverse mask ,in, and are used to represent the context information of the target area and the background area respectively; and After splicing in the channel dimension, it is input into the 1×1 convolution and Sigmoid activation function to generate a set of Pixel weight matrix with the same resolution and number of channels , used for The features are recalibrated to obtain the self-guided correction features , the calculation formula is:

[0104] ;

[0105] Where, Indicates the mask Take the inverse, that is, from the matrix Subtract ; The output of the feature adaptive fusion module is composed of the aligned and corrected high-level features and the low-level features after self-guided correction After weighted fusion and feature integration through 3×3 convolution, the formula is:

[0106] ;

[0107] Where, is the output of the feature adaptive fusion module.

[0108] Specifically, the feature adaptive fusion module receives low-level and high-level feature maps of adjacent scales. The high-level feature map is upsampled using bilinear interpolation to maintain the same spatial resolution as the low-level feature map. A 3×3 convolution is performed on the upsampled high-level feature map to adjust the number of channels to match that of the low-level feature map. The high-level feature map is then element-wise added to the low-level feature map to achieve preliminary feature fusion. To reduce spatial position information errors that may arise during feature fusion, 3×3 dilated convolutions with dilation rates of 1, 3, and 5 are applied to the fused features, resulting in three sets of feature representations with different receptive fields. These representations are then concatenated along the channel dimension, and then restored using 1×1 convolutions.

[0109] For low-level feature maps First, a 1×1 convolution combined with a Sigmoid activation function is used to compress the features and obtain the pixel-level guidance mask. (all its elements have the value 1) minus The inverse mask is obtained by concatenating the mask and the inverse mask in the channel dimension and inputting them into the 1×1 convolution and Sigmoid activation function to generate a set of Pixel weight matrix with the same resolution and number of channels , used for The features are recalibrated to obtain the self-guided correction features; finally, the output of the feature adaptive fusion module is composed of the aligned and corrected high-level features and the low-level features after self-guided correction The weighted fusion is performed and the features are integrated through 3×3 convolution.

[0110] Through the above steps, the complementary enhancement between features of different scales is effectively achieved, and the detection performance of target areas in remote sensing images is improved.

[0111] In one embodiment, after step S130 and before step S140, the following step is further included: converting the binary cross entropy loss With Dice loss Perform weighted fusion and construct a joint loss function. The formula is:

[0112] ;

[0113] ;

[0114] ;

[0115] Where, and Respectively and The maximum value of and Represents the position in the image The predicted value and the true label at Represents the cross entropy loss function, symbol express norm; the feature adaptive fusion modules are stacked in a step-by-step manner to produce three output results with the same resolution as the lowest layer feature map; the loss values of the three output results are calculated through the joint loss function, and the three calculated loss values are assigned different weight coefficients and weighted summed to obtain the final loss value. The formula is:

[0116] ;

[0117] ;

[0118] Where, 、 and They represent the loss values calculated from the output results of the three lowest-level feature adaptive fusion modules, 、 and are the weight coefficients assigned to the corresponding loss items, is the final loss value.

[0119] Specifically, in order to effectively evaluate model performance and improve detection accuracy, the predicted changes in the model output are compared with the ground truth labels during the model training phase, the loss value of each output is calculated, and the model parameters are optimized accordingly. Since open-pit mining areas usually account for a small proportion in remote sensing images and there is a significant class imbalance problem, if only the binary cross entropy loss function is used, it is easy to cause loss function oscillation and unstable training process. To this end, the present invention introduces the Dice loss function to enhance the sensitivity and robustness of the model to small target areas.

[0120] Furthermore, to balance classification accuracy and small object detection capabilities, a weighted fusion of binary cross entropy loss and Dice loss is performed to construct a joint loss function. Considering that the model structure contains multiple feature adaptive fusion modules, independent losses are calculated for the feature maps output by each module during training. These losses are then weighted and summed according to the set weight coefficients to form the overall loss function, thereby achieving deep supervision and collaborative optimization of feature learning at all stages of the model.

[0121] Because the feature adaptive fusion modules are stacked in a stepped fashion, three outputs with the same resolution as the lowest-level feature map are generated during the model's forward propagation. To further improve convergence speed and stability during model training, a joint loss function is used to independently calculate the corresponding loss values for these three outputs. The final total loss function assigns different weights to the three loss values and then takes a weighted sum to obtain a comprehensive optimization objective, which facilitates further model optimization based on the loss function.

[0122] Step S140: inputting the remote sensing image data to be tested into the remote sensing image detection model to obtain the detection result of the open-pit mining area.

[0123] Specifically, after the remote sensing image detection model training is completed, the remote sensing image data to be tested is input into the remote sensing image detection model to generate a binary result map, and the remote sensing open-pit mine detection results are obtained based on the binary result map. The detection results are analyzed and interpreted, and the mining area in the remote sensing image is identified, thereby obtaining information on resource exploration, mineral management, and environmental monitoring.

[0124] Through the trained remote sensing image detection model, the remote sensing image data to be tested is inferred, and the output results after multi-scale fusion are combined to obtain a refined mining area detection map. Then, based on the mining area detection results in all remote sensing images, the spatial distribution information of the open-pit mining area and the structural feature analysis are realized.

[0125] In this embodiment, a multi-phase remote sensing image dataset is constructed by acquiring multi-source high-resolution remote sensing image data at multiple time points and covering the entire range of the target open-pit mining area, and the dataset is preprocessed with atmospheric correction, radiation calibration, color correction, image registration and cropping operations to obtain time series remote sensing image data to improve image quality and time series consistency; a deep learning algorithm is used to construct an open-pit mine area detection model based on a dual-branch structure and a feature fusion mechanism, and the model is trained with time series remote sensing image data to obtain a remote sensing image detection model, wherein the dual-branch structure comprises a feature extraction branch and a feature generation branch, and the feature fusion mechanism comprises a cross-attention fusion module and a feature adaptive fusion model, thereby effectively realizing the dynamic aggregation and expression enhancement of features of different scales and different semantic levels, and inputting the remote sensing image data to be tested into the remote sensing image detection model, and the model is trained with time series remote sensing image data. Through the dual-branch structure, the mining area information in the remote sensing image is modeled from the two perspectives of spatial details and semantic structure, and two feature fusion modules are used to realize information interaction and supplement between branches, further improving the detection accuracy and model generalization ability, and obtaining the detection results of the open-pit mining area. Combined with the dual-branch structure and feature fusion mechanism, it can deeply mine the structural information and detail features in the remote sensing image, and strengthen the feature fusion expression, thereby improving the model's recognition ability for complex landforms, ensuring the recognition accuracy of small target details and mining area boundaries in low-resolution images, thereby improving the accuracy and robustness of remote sensing image detection, and can be applied to target detection tasks in multi-scale and diversified remote sensing scenarios. In addition, the present invention has a low parameter amount and can achieve efficient calculation, thereby ensuring both detection accuracy and calculation efficiency, and improving the deployment efficiency and practicality of the model.

[0126] like Figure 2 Figure 2 shows the performance of the algorithm of the present invention compared with existing algorithms such as GT (Gradient Boosting Tree), U-Net, SegNet, DU SegNet, GIPNet, ASNet, Easy-Net, and DASFNet in detecting targets with building interference, abnormal exposure, small targets, and complex edge structures. Clearly, the present invention is capable of high-precision detection of targets in environments with interference and abnormal exposure, and is also suitable for detecting small targets and targets with complex edge structures.

[0127] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0128] Obviously, those skilled in the art will appreciate that the various modules or steps of the present invention described above can be implemented using a general-purpose computing device. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Alternatively, they can be implemented using program code executable by the computing device, which can then be stored in a computer storage medium (ROM / RAM, magnetic disk, optical disk) and executed by the computing device. In some cases, the steps shown or described herein can be performed in a different order than that shown, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Therefore, the present invention is not limited to any particular combination of hardware and software.

[0129] The above content is a further detailed description of the present invention in conjunction with specific embodiments, and the specific implementation of the present invention cannot be considered to be limited to these descriptions. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should be considered to fall within the scope of protection of the present invention.

Claims

1. A remote sensing open-pit mine detection method based on a dual-branch structure and feature fusion mechanism, characterized in that: The following steps are involved: Acquire multi-source high-resolution remote sensing image data of open-pit mining areas; Preprocessing the remote sensing image data to obtain time-series remote sensing image data, wherein the preprocessing includes atmospheric correction, radiation calibration, color correction, image registration, and cropping operations; A deep learning algorithm is used to construct an open-pit mine area detection model based on a dual-branch structure and a feature fusion mechanism, and the model is trained using the time-series remote sensing image data to obtain a remote sensing image detection model. The dual-branch structure includes a feature extraction branch and a feature generation branch, and the feature fusion mechanism includes a cross-attention fusion module and a feature adaptive fusion module. The feature extraction branch uses a lightweight MobileNet V2 as the backbone network, and removes the final pooling layer and fully connected layer to extract multi-scale preliminary features from the remote sensing image data; The feature generation branch is composed of multiple dense generation modules, which are used to generate multi-scale fine-grained features from the remote sensing image data to enhance the learning of mining area edges and texture details; The cross attention fusion module is used to supplement the multi-scale fine-grained features into the multi-scale preliminary features; The feature adaptive fusion module adopts a ladder structure for stacking, and performs fusion of adjacent scale features and interaction of non-adjacent scale features during the decoding process to fully capture the detailed information of the mining area target; the remote sensing image data to be tested is input into the remote sensing image detection model to obtain the detection results of the open-pit mining area.

2. The remote sensing open-pit mine area detection method based on a dual-branch structure and feature fusion mechanism according to claim 1 is characterized in that: The MobileNet V2 network includes an input layer, an initial convolutional layer, multiple inverted residual structure modules, a global average pooling layer, and a fully connected layer; Wherein, the input layer is used to receive pre-processed time series remote sensing image data; The initial convolution layer uses a convolution kernel of size 3×3 and a step size of 2 to perform preliminary feature extraction on the input image; The inverted residual structure module includes an expansion layer, a depthwise separable convolution layer, and a projection layer, wherein the expansion layer is a 1×1 convolution operation for expanding the channel dimension of the input features; the depthwise separable convolution layer is composed of a 3×3 depthwise convolution and a 1×1 pointwise convolution, wherein the depthwise convolution is used to perform spatial convolution operations on each channel independently, and the pointwise convolution is used to integrate information between channels; the projection layer uses a 1×1 convolution to compress the high-dimensional features back to the original low-dimensional space and fuse them with the input features through jump connections; Multiple inverted residual structure modules are stacked with different numbers of channels and resolutions to construct a deep feature extraction network; The global average pooling layer is used to compress the spatial dimension of the feature map; The fully connected layer is used to map global features to classification results or detection outputs.

3. The remote sensing open-pit mine area detection method based on a dual-branch structure and feature fusion mechanism according to claim 1 is characterized in that: The dense generation module includes three residual dense blocks, each of which includes five 3×3 convolutional layers connected in sequence, and Leaky ReLU is used as the activation function after each convolution operation, and does not include a batch normalization layer; The convolutional layers are connected by jump connections through channel dimension splicing, the residual dense blocks are fused by element-by-element addition, and the residual scaling coefficient is introduced to scale the fusion result; The structural end of the dense generation module uses a 3×3 convolutional layer with a stride of 2 for downsampling to keep the size of the feature map consistent with the feature extraction branch; Among them, the formula of the dense generation module is: ; Where, Represents a convolution operation with a kernel size of 3 and a stride of 1; Indicates the first The output result of the convolution operation; Indicates that in the current residual dense block, the features of the previous layers are concatenated in the channel dimension before the i-th convolution; Represent the output results of the three residual dense blocks respectively; Represents the Leaky ReLU activation function.

4. The remote sensing open-pit mine area detection method based on a dual-branch structure and feature fusion mechanism according to claim 2 is characterized in that: The cross attention fusion module includes a spatial attention block and a channel attention block; Through the spatial attention block and the channel attention block, the multi-scale preliminary features and the multi-scale fine-grained features are subjected to attention operations in the spatial dimension and the channel dimension respectively, to obtain a channel attention sequence and a spatial attention feature map, and a cross matrix multiplication operation is performed to obtain two weight matrices; Based on the two weight matrices, a re-weighting operation of element-by-element multiplication is performed on the multi-scale preliminary features and the multi-scale fine-grained features respectively; Perform a feature fusion operation on the re-weighted features by adding them element by element to obtain the fused features; The fusion feature is adjusted through the depthwise separable convolutional layer to obtain an output feature after fusing the two branch features.

5. The remote sensing open-pit mine detection method based on a dual-branch structure and feature fusion mechanism according to claim 4 is characterized in that: The channel attention block is used to perform global maximum pooling and global average pooling on the input feature map in the spatial dimension to obtain two sets of different channel information sequences; The two sets of channel information sequences are input into a multi-layer perceptron through the channel attention block to perform feature reconstruction and extract high-level channel features; The two sets of channel features after feature reconstruction are fused through element-level addition to obtain the output channel attention sequence containing rich channel information of the input feature map; Among them, the formula of the channel attention block is: ; Where, and Represent global average pooling and global maximum pooling operations respectively, MLP represents multi-layer perceptron, represents the input features, Represents the output sequence of the channel attention module.

6. The remote sensing open-pit mine detection method based on a dual-branch structure and feature fusion mechanism according to claim 4 is characterized in that: The spatial attention block is used to enhance the semantic representation capability of the feature map; The spatial attention block performs global maximum pooling and global average pooling operations on the channel dimension of the input feature map, and combines 1×1 convolution to compress the channel dimension of the input feature map to generate three spatial information matrices; The three spatial information matrices are spliced in the channel dimension and compressed through convolution operation to obtain the output spatial information matrix , used to guide the feature enhancement of important spatial regions, the formula is: ; ; ; ; Where, and Represent the maximum pooling operation and average pooling operation in the channel dimension respectively; Represents the output spatial information matrix of the channel attention module; Represents a point convolution operation with a convolution kernel size of 1; Indicates that the concatenation operation is performed on the channel dimension.

7. The remote sensing open-pit mine detection method based on a dual-branch structure and feature fusion mechanism according to claim 1 is characterized in that: The feature adaptive fusion module is used to receive two sets of feature maps at adjacent scales, namely low-level feature maps With high-level feature maps ; Performing bilinear interpolation upsampling processing on the high-level feature map through the feature adaptation module so that its spatial resolution is consistent with that of the low-level feature map; Perform 3×3 convolution on the upsampled high-level feature map to make it consistent with the number of channels of the low-level feature map and Perform element-level addition operation to perform preliminary feature fusion. The formula is: ; Where, Represents a bilinear interpolation upsampling operation; Fusion features 3×3 dilated convolution operations with dilation rates of 1, 3, and 5 are used respectively to obtain feature representations under three different receptive fields, which are then concatenated in the channel dimension. 3×3 convolution is used to restore the channel dimension. The process is as follows: ; ; Where, The expansion rate is Convolution operation; For low-level feature maps , feature compression is performed through a 1×1 convolution combined with a Sigmoid activation function to obtain a pixel-level guidance mask , whose expression is: ; By setting all elements of the matrix to 1 minus Get the reverse mask ,in, and Used to represent the context information of the target area and background area respectively; Will and After splicing in the channel dimension, it is input into the 1×1 convolution and Sigmoid activation function to generate a set of Pixel weight matrix with the same resolution and number of channels , used for The features are recalibrated to obtain the self-guided correction features , the calculation formula is: ; Where, Indicates the mask Take the inverse, that is, from the matrix Subtract ; The output of the feature adaptive fusion module is composed of the aligned and corrected high-level features and the low-level features after self-guided correction After weighted fusion and feature integration through 3×3 convolution, the formula is: ; Where, is the output of the feature adaptive fusion module.

8. The remote sensing open-pit mine area detection method based on a dual-branch structure and feature fusion mechanism according to claim 1 is characterized in that: After the open-pit mine area detection model based on the dual-branch structure and feature fusion mechanism is constructed using the deep learning algorithm and trained with the time series remote sensing image data to obtain the remote sensing image detection model, and before the remote sensing image data to be tested is input into the remote sensing image detection model to obtain the detection result of the open-pit mine mining area, the method further includes: Binary cross entropy loss With Dice loss Perform weighted fusion and construct a joint loss function. The formula is: ; ; ; Where, and Respectively and The maximum value of and Represents the position in the image The predicted value and the true label at Represents the cross entropy loss function, symbol express norm; The feature adaptive fusion modules are stacked in a step-by-step manner to generate three output results that are consistent with the resolution of the lowest layer feature map; The loss values of the three output results are calculated by the joint loss function, and the three calculated loss values are assigned different weight coefficients and weighted summed to obtain the final loss value. The formula is: ; ; Where, 、 and They represent the loss values calculated from the output results of the three lowest-level feature adaptive fusion modules, 、 and are the weight coefficients assigned to the corresponding loss items, is the final loss value.

Citation Information

Patent Citations

  • Remote sensing image change detection method based on twinborn multi-scale cross attention

    CN117975267A

  • Land change detection method combining matrix decomposition and adaptive propagation, and system

    WO2025111921A1