Remote sensing strip mine area detection method based on double-branch structure and feature fusion mechanism
Through the deep learning method of the dual-branch structure and feature fusion mechanism, the problem of low recognition accuracy of open-pit mining areas and small targets in remote sensing images is solved, and efficient and accurate mining area detection is achieved, which is suitable for a variety of remote sensing scenarios.
Patent Information
- Application Number
- CN202510757188.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-09
AI Technical Summary
The prior art is difficult to accurately identify the boundary and land features of open-pit mining areas in remote sensing images, especially in low-resolution images, small target details recognition accuracy, and traditional methods have high computational complexity, making it difficult to apply in resource-constrained environments.
Deep learning methods based on dual-branch structure and feature fusion mechanism are adopted, including feature extraction branches and feature generation branches, combined with cross attention fusion module and feature adaptive fusion module, and train the model through multi-source high-resolution remote sensing image data to improve feature extraction and fusion capabilities.
It improves the recognition accuracy of mining area boundaries and small target details in low-resolution images, reduces the computational complexity, enhances the robustness and adaptability of the model, and is suitable for multi-scale and diverse remote sensing scenarios.
Smart Images

Figure CN120279437A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and particularly to a remote sensing open-pit mining area detection method based on a dual-branch structure and a feature fusion mechanism. Background Art
[0002] Traditional mineral resource supervision modes mainly rely on manual inspections and on-site law enforcement. This method is not only time-consuming and laborious, but also greatly affected by factors such as terrain and climate, making it difficult to achieve efficient and accurate monitoring of large-scale areas. With the rapid development of remote sensing technology and the continuous improvement of aerospace detection capabilities, the ability to obtain high-resolution remote sensing images has been significantly enhanced, greatly promoting the wide application of remote sensing data in multiple fields, such as land resource management, urban planning, crop management, environmental protection, and natural disaster monitoring. At the same time, it has also brought new opportunities for dynamic monitoring of mineral resources. Remote sensing images have the characteristics of large-scale coverage and periodic updates, and can continuously obtain surface information of mining areas, providing accurate and efficient data support for dynamic monitoring of mining activities. At present, using high-resolution remote sensing images for automatic detection and identification of open-pit mining areas has become an important technical means for mineral resource management.
[0003] However, due to the special imaging conditions of remote sensing images and the complex geographical environment of open-pit mining areas, current remote sensing detection of open-pit mining areas still faces many challenges. First, due to the limitations of factors such as imaging equipment and shooting height on the resolution of remote sensing images, many images still have low-resolution problems, resulting in relatively blurred surface texture information of mining areas and making it difficult to clearly extract the boundaries and ground object features of mining areas. Second, the lighting conditions and imaging angles of remote sensing images vary greatly, easily causing uneven color distribution, resulting in a low contrast between mining areas and the surrounding environment, thus increasing the difficulty of target recognition. In addition, open-pit mining areas usually have relatively complex topographical and geomorphic features, and there are large differences in the ground object attributes of different mineral types. It is difficult to accurately model and describe their features such as boundaries, shapes, and spatial distributions. These factors seriously affect the detection accuracy of existing methods. Therefore, how to effectively overcome problems such as low resolution, blurred texture information, and uneven color of remote sensing images, accurately analyze the structural features of mining areas (such as boundaries, shapes, spatial distributions, etc.), and improve the accuracy of target recognition is an important problem that urgently needs to be solved in the current research on remote sensing open-pit mining area detection.
[0004] In the early stage, the identification of open-pit mining areas mainly relied on manual visual interpretation and manual annotation methods. However, this method not only required a large amount of manpower and time, but also highly depended on expert experience, with problems such as strong subjectivity, low efficiency, and unstable accuracy. In addition, affected by factors such as the quality of remote sensing images, weather conditions, and lighting changes, it was often difficult to accurately define the boundaries of mining areas through manual interpretation, resulting in large errors in detection results. With the development of machine learning technology, researchers gradually explored applying machine learning methods to remote sensing mining area detection tasks and continuously optimized algorithms to improve detection accuracy. In machine learning-based methods, mining area detection is usually regarded as a pixel-level classification task of remote sensing images, that is, each pixel is classified into mining area or non-mining area categories, and automatic recognition is achieved through classification algorithms in machine learning. Although compared with traditional manual interpretation methods, machine learning methods have improved detection efficiency and reduced subjective errors to a certain extent, early machine learning methods still faced many challenges. For example, in the face of a large amount of remote sensing data, how to efficiently process and extract key features remained a difficult problem. In addition, the terrain and landforms of mining areas were complex, and there were significant differences in spectral features in different bands. A single machine learning model was difficult to accurately depict the boundaries and spatial distribution relationships of mining areas, resulting in limited detection accuracy.
[0005] In recent years, with the rapid development of deep learning, deep learning-based methods have achieved remarkable success in the field of image processing. With its powerful non-linear mapping ability and end-to-end learning mechanism, it has greatly improved the performance of object detection, segmentation, and recognition tasks. Compared with traditional methods that rely on manually designed features, deep learning can automatically learn effective features through large-scale data training, avoiding the subjectivity of feature selection, and at the same time improving the adaptability and generalization ability of the model. By constructing a multi-layer neural network, the deep learning model can extract semantic information from low-level to high-level layer by layer, so as to more accurately understand the target structure in the image and its relationship with the background. In addition, the introduction of the attention mechanism further enhances the model's ability to model global information, making it perform better when dealing with complex environments and long-range dependence relationships. However, due to the complex surface environment of mining areas, the target boundaries often show a gradual transition, making it difficult for the detection model to accurately distinguish mining areas from surrounding ground objects, resulting in prominent problems of false detection and missed detection. Secondly, limited by the acquisition conditions of remote sensing images, the resolution of many optical remote sensing data is low, and each pixel may contain multiple types of ground objects, making it more difficult to effectively extract texture and structure information. In addition, lighting changes, shadow interference, and uneven color distribution of different landforms in mining areas limit the generalization ability of traditional deep learning methods in different scenarios. Currently, most deep learning methods mainly rely on deepening the network structure or increasing the number of trainable parameters to improve detection performance, but this often leads to a significant increase in computational complexity and storage requirements, restricting the practical application of the model in resource-constrained environments.
[0006] Therefore, for the remote sensing detection task of open-pit mining areas, there is an urgent need for a lightweight remote sensing open-pit mining area detection method that can improve the recognition accuracy of small target details and mining area boundaries in low-resolution images, and balance computational efficiency and detection accuracy. Summary of the Invention
[0007] Based on this, it is necessary to provide a remote sensing open-pit mining area detection method based on a double-branch structure and a feature fusion mechanism for the above technical problems.
[0008] A remote sensing open-pit mining area detection method based on a double-branch structure and a feature fusion mechanism includes the following steps: obtaining multi-source high-resolution remote sensing image data of an open-pit mining area; preprocessing the remote sensing image data to obtain time-series remote sensing image data, and the preprocessing includes atmospheric correction, radiometric calibration, color correction, image registration, and cropping operations; constructing an open-pit mining area detection model based on a double-branch structure and a feature fusion mechanism using a deep learning algorithm, and training it with the time-series remote sensing image data to obtain a remote sensing image detection model. The double-branch structure includes a feature extraction branch and a feature generation branch, and the feature fusion mechanism includes a cross-attention fusion module and a feature adaptive fusion module; among them, the feature extraction branch uses the lightweight MobileNet V2 as the backbone network and removes the final pooling layer and fully connected layer to extract multi-scale preliminary features from the remote sensing image data; the feature generation branch is composed of multiple dense generation modules to generate multi-scale fine-grained features from the remote sensing image data to enhance the learning of the mining area edge and texture details; the cross-attention fusion module is used to supplement the multi-scale fine-grained features to the multi-scale preliminary features; the feature adaptive fusion module is stacked in a ladder structure to perform the fusion of adjacent-scale features and the feature interaction of non-adjacent scales during the decoding process to comprehensively capture the detailed information of the mining area target; inputting the remote sensing image data to be tested into the remote sensing image detection model to obtain the detection result of the open-pit mining area.
[0009] In one embodiment, the MobileNet V2 network includes an input layer, an initial convolutional layer, multiple inverted residual structure modules, a global average pooling layer, and a fully connected layer. Among them, the input layer is used to receive the preprocessed temporal remote sensing image data. The initial convolutional layer uses a convolutional kernel of size 3×3 and a stride of 2, and is used to perform preliminary feature extraction on the input image. The inverted residual structure module includes an expansion layer, a depthwise separable convolutional layer, and a projection layer. Among them, the expansion layer is a 1×1 convolutional operation, which is used to expand the input features in the channel dimension. The depthwise separable convolutional layer consists of a 3×3 depthwise convolution and a 1×1 pointwise convolution. Among them, the depthwise convolution is used to independently perform spatial convolution operations on each channel, and the pointwise convolution is used to perform information integration between channels. The projection layer uses a 1×1 convolution to compress the high-dimensional features back to the original low-dimensional space and fuse them with the input features through skip connections. Multiple inverted residual structure modules are stacked with different channel numbers and resolutions to construct a deep feature extraction network. The global average pooling layer is used to compress the feature map in the spatial dimension. The fully connected layer is used to map the global features to the classification results or detection outputs.
[0010] In one embodiment, the dense generation module includes three residual dense blocks. Each residual dense block includes five sequentially connected 3×3 convolutional layers. After each convolutional operation, Leaky ReLU is used as the activation function, and there is no batch normalization layer. The convolutional layers are connected by skip connections through channel dimension concatenation. The residual dense blocks are fused by element-wise addition, and a residual scaling coefficient is introduced to scale the fusion result. The end of the structure of the dense generation module uses a 3×3 convolutional layer with a stride of 2 for downsampling, so that the size of the feature map is consistent with the feature extraction branch. Among them, the formula of the dense generation module is: ; In the formula, represents a convolutional operation with a convolutional kernel size of 3 and a stride of 1. represents the output result of the th convolutional operation in the current residual dense block. represents that in the current residual dense block, before the i-th convolution, the features of the previous layers are concatenated in the channel dimension. respectively represent the output results of the three residual dense blocks. represents the Leaky ReLU activation function.
[0011] In one embodiment, the cross-attention fusion module includes a spatial attention block and a channel attention block; through the spatial attention block and the channel attention block, attention operations are respectively performed on the multi-scale preliminary features and the multi-scale fine-grained features in the spatial dimension and the channel dimension to obtain a channel attention sequence and a spatial attention feature map, and a cross-matrix multiplication operation is performed to obtain two weight matrices; based on the two weight matrices, an element-wise multiplication reweighting operation is respectively performed on the multi-scale preliminary features and the multi-scale fine-grained features; a feature fusion operation of element-wise addition is performed on the reweighted features to obtain fused features; the fused features are adjusted through the depthwise separable convolution layer to obtain output features after fusing the features of the two branches.
[0012] In one embodiment, the channel attention block is used to perform global max pooling and global average pooling respectively on the input feature map in the spatial dimension to obtain two groups of different channel information sequences; the two groups of channel information sequences are input into a multi-layer perceptron by the channel attention block for feature reconstruction to extract high-level channel features; the two groups of channel features after feature reconstruction are fused through element-wise addition to obtain an output channel attention sequence containing rich channel information of the input feature map; where the formula of the channel attention block is: ; In the formula, and respectively represent the global average pooling and global max pooling operations, MLP represents the multi-layer perceptron, represents the input feature, represents the output sequence of the channel attention module.
[0013] In one embodiment, the spatial attention block is used to enhance the semantic representation ability of the feature map; through the spatial attention block, global max pooling and global average pooling operations are respectively performed on the channel dimension of the input feature map, and the channel dimension of the input feature map is compressed by combining 1×1 convolution to generate three spatial information matrices; the three spatial information matrices are concatenated in the channel dimension and compressed through a convolution operation to obtain the output spatial information matrix , which is used to guide the feature enhancement of important spatial regions, and the formula is: ; ; ; ; In the formula, and respectively represent the max pooling operation and the average pooling operation in the channel dimension; Denote the output spatial information matrix of the channel attention module; Denote the point convolution operation with a convolution kernel size of 1; Denote the concatenation operation in the channel dimension.
[0014] In one embodiment, the feature adaptive fusion module is used to receive two sets of feature maps at adjacent scales, namely the low-level feature map and the high-level feature map ; perform bilinear interpolation upsampling processing on the high-level feature map through the feature adaptive module to make its spatial resolution consistent with that of the low-level feature map; perform 3×3 convolution on the upsampled high-level feature map to make its number of channels consistent with that of the low-level feature map, and perform an element-wise addition operation with the low-level feature map to perform preliminary feature fusion. The formula is: ; In the formula, Denote the bilinear interpolation upsampling operation; for the fused feature perform 3×3 dilated convolution operations with dilation rates of 1, 3, and 5 respectively to obtain feature representations under three different receptive fields, and perform concatenation in the channel dimension, and use 3×3 convolution to restore the channel dimension. The process is: ; ; In the formula, Denote the convolution operation with a dilation rate of ; for the low-level feature map , perform feature compression through a 1×1 convolution combined with a Sigmoid activation function to obtain a pixel-level guidance mask , and its expression is: ; By subtracting the matrix with all element values of 1 from , obtain the anti-mask , where and are used to represent the context information of the target region and the background region respectively; after concatenating and in the channel dimension, input them into a 1×1 convolution and a Sigmoid activation function to generate a set of pixel weight matrices with the same resolution and number of channels as , which are used to recalibrate the features of to obtain the self-guided corrected features . The calculation formula is: ; In the formula, represents taking the inverse of the mask , that is, subtracting from the matrix ; the output of the feature adaptive fusion module is obtained by weighted fusion of the high-level features after alignment and correction and the low-level features after self-guided correction, and then integrating the features through 3×3 convolution. The formula is: ; In the formula, is the output of the feature adaptive fusion module.
[0015] In one embodiment, after constructing an open-pit mining area detection model based on a double-branch structure and a feature fusion mechanism using a deep learning algorithm and training it with the time-series remote sensing image data to obtain a remote sensing image detection model, before inputting the remote sensing image data to be measured into the remote sensing image detection model to obtain the detection result of the open-pit mining area, it further includes: weighting and fusing binary cross-entropy loss and Dice loss to construct a joint loss function. The formula is: ; ; ; In the formula, and respectively represent the maximum values of and , and respectively represent the predicted value and the true label at the position in the image; the function represents the cross-entropy loss function, and the symbol represents norm; stacking the feature adaptive fusion module step by step to generate three output results with the same resolution as the lowest-level feature map; calculating the loss values of the three output results through the joint loss function, and performing weighted summation after assigning different weight coefficients to the calculated three loss values to obtain the final loss value. The formula is: ; ; In the formula, , and respectively represent the loss values calculated from the output results of the three lowest-level feature adaptive fusion modules, , and are the weight coefficients assigned to the corresponding loss terms, is the final loss value.
[0016] Compared with the prior art, the advantages and beneficial effects of the present invention are as follows: by acquiring multi-source high-resolution remote sensing image data covering the entire open-pit mining area from different satellites, and performing preprocessing such as atmospheric correction, radiation calibration, color correction, image registration and cropping operations, time-series remote sensing image data is obtained, and the contrast of the image is improved, making it more conducive to model detection; a deep learning algorithm is used to construct an open-pit mining area detection model based on a dual-branch structure and a feature fusion mechanism, and the remote sensing image detection model is obtained by training with time-series remote sensing image data, the dual-branch structure includes a feature extraction branch and a feature generation analysis, and the feature fusion mechanism includes a cross-attention fusion module and a feature adaptive fusion module, The trained remote sensing image detection model is used to detect the remote sensing image data to be tested, and the detection results of the open-pit mining area are obtained; combined with the dual-branch structure and feature fusion mechanism, it is possible to deeply mine the structural information and detail features in the remote sensing image, and strengthen the feature fusion expression, thereby improving the model's recognition ability for complex objects, ensuring the recognition accuracy of small target details and mining area boundaries in low-resolution images, thereby improving the accuracy and robustness of remote sensing image detection, and can be applicable to target detection tasks in multi-scale and diversified remote sensing scenarios. In addition, the present invention has a low number of parameters and can achieve efficient calculation, thereby ensuring both detection accuracy and calculation efficiency, thereby improving the deployment efficiency and practicability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 A schematic diagram of a flow chart of a remote sensing open-pit mine area detection method based on a dual-branch structure and a feature fusion mechanism in one embodiment; Figure 2 A schematic diagram of the architecture of a remote sensing image detection model in one embodiment; Figure 3 is a schematic diagram of the structure of a cross attention fusion module in one embodiment; Figure 4 is a schematic diagram of the structure of a feature adaptive fusion module in an embodiment; Figure 5 is a schematic diagram of the structure of a dense generation module in one embodiment; Figure 6 The diagram is a visual comparison diagram of the present invention and other existing methods on a data set in one embodiment. DETAILED DESCRIPTION
[0018] Before describing the specific embodiments of the present invention, the overall concept of the present invention is described as follows: The present invention is mainly developed based on the remote sensing data processing process. When facing low-resolution image detection, traditional remote sensing data processing methods have poor recognition performance for small target details and mining area boundaries, low detection accuracy, or complex model structure, which is not conducive to practical application.
[0019] Therefore, the present invention proposes a remote sensing open-pit mine area detection method based on a dual-branch structure and a feature fusion mechanism, obtains multi-source high-resolution remote sensing image data of the open-pit mine sampling area, forms a remote sensing image data set, and performs preprocessing such as atmospheric correction, radiation calibration, color correction, image registration and cropping operations on the remote sensing image data to improve image quality and subsequent modeling effects; a deep learning algorithm is used to construct an open-pit mine area detection model based on a dual-branch structure and a feature fusion mechanism, and the remote sensing image detection model is obtained by training with time-series remote sensing image data. The dual branches include a feature extraction branch and a feature generation branch, and the feature fusion mechanism of the cross-attention fusion module and the feature adaptive fusion module are introduced to achieve efficient feature interaction and fusion between the dual branches, thereby enhancing The ability to extract target features of mining areas in complex remote sensing images is improved by using the trained remote sensing image detection model to detect and process the remote sensing image data to obtain accurate identification results of open-pit mining areas. Combined with the dual-branch structure and feature fusion mechanism, it can deeply mine the structural information and detail features in the remote sensing images, strengthen the feature fusion expression, and enhance the model's ability to recognize complex objects. It ensures the recognition accuracy of small target details and mining area boundaries in low-resolution images, thereby improving the accuracy and robustness of remote sensing image detection, and can be applied to target detection tasks in multi-scale and diversified remote sensing scenarios. In addition, the present invention has a low number of parameters and can achieve efficient calculation, thereby ensuring detection accuracy while taking into account calculation efficiency, thereby improving the deployment efficiency and practicality of the model.
[0020] After introducing the overall concept of the present invention, in order to make the purpose, technical solution and advantages of the present invention more clear, the present invention is further described in detail by specific implementation methods in combination with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0021] In one embodiment, Figure 1 As shown, a remote sensing open-pit mine area detection method based on a dual-branch structure and a feature fusion mechanism is provided, comprising the following steps: Step S110, obtaining multi-source high-resolution remote sensing image data of the open-pit mine mining area.
[0022] Specifically, multi-source high-resolution remote sensing images covering the open-pit mine mining area are obtained through various types of remote sensing satellites. The remote sensing image data covers the data collected by multiple satellite platforms, has a high spatial resolution and imaging frequency, and can adapt to the imaging requirements under different terrain conditions. The collected remote sensing image data covers various types of open-pit mining areas (such as sandstone, limestone, shale, etc.), ensuring that the detection model can be applied to the identification and extraction of different mineral species, and has good generality and adaptability.
[0023] Step S120: Preprocess the remote sensing image data to obtain time-series remote sensing image data. The preprocessing includes atmospheric correction, radiometric calibration, color correction, image registration, and cropping operations.
[0024] Specifically, systematically preprocess the obtained satellite remote sensing image data. First, perform atmospheric correction to remove the effects of atmospheric scattering and absorption, etc.; then perform radiometric calibration to unify the physical response values of imaging data from different sensors; subsequently, carry out color correction to improve the color consistency and visual effect of the image; precisely geometrically register the images from different satellite platforms or imaging times to ensure the consistency of spatial positions; finally, crop the image according to the boundary of the open-pit mining area to extract the target area image segment. Through the above preprocessing operations, time-series remote sensing image data is obtained for subsequent model training, significantly improving the quality and consistency of the remote sensing data, and laying a foundation for the accurate detection and feature extraction of subsequent models.
[0025] Step S130: Use a deep learning algorithm to construct an open-pit mining area detection model based on a double-branch structure and a feature fusion mechanism, and train it with time-series remote sensing image data to obtain a remote sensing image detection model. The double-branch structure includes a feature extraction branch and a feature generation branch, and the feature fusion mechanism includes a cross-attention fusion module and a feature adaptive fusion module.
[0026] Among them, the feature extraction branch uses the lightweight MobileNet V2 as the backbone network and removes the final pooling layer and fully connected layer to extract multi-scale preliminary features from the remote sensing image data; the feature generation branch is composed of multiple dense generation modules to generate multi-scale fine-grained features from the remote sensing image data, enhancing the learning of the mining area edge and texture details; the cross-attention fusion module is used to supplement the multi-scale fine-grained features to the multi-scale preliminary features; the feature adaptive fusion module is stacked in a ladder structure to perform the fusion of adjacent-scale features and the feature interaction of non-adjacent scales during the decoding process, comprehensively capturing the detailed information of the mining area target.
[0027] Specifically, a deep learning algorithm is used to construct an open-pit mining area detection model based on a dual-branch structure and a feature fusion mechanism, and the preprocessed time-series remote sensing image data is used for training. After the training is completed, a remote sensing image detection model is obtained, as Figure 2 shown, which is used for the detection of open-pit mining areas and is suitable for the precise detection of open-pit mines in complex environments. Among them, the dual-branch structure includes a feature extraction branch and a feature generation branch, and the feature fusion mechanism includes a cross-attention fusion module (Cross-Attention Fusion Module, CAFM) and a feature adaptive fusion module (Feature AdaptiveFusion Module, FAFM).
[0028] As Figure 2 shown, it is a schematic diagram of the structure of the remote sensing image detection model. The input image is processed through the feature extraction branch and the feature generation branch respectively. The feature extraction branch uses stage1-stage5 for feature extraction and inputs it to the cross-attention fusion module for feature fusion, so that the model can ensure the global semantic understanding ability while fully obtaining local detailed features; at the same time, the dense generation module (Dense GenerationModule, DGM) of the feature generation branch is used for efficient interaction between features of different levels, and the fused features are input to the stepped feature adaptive module for the fusion of adjacent scale features and the interaction of non-adjacent scale features. Finally, the detection result is output, and the output detection result can be feedback supervised through the ground truth to ensure the accuracy of the model, so that the model can comprehensively capture the detailed information of the mining area target. The detection accuracy of the remote sensing image detection model obtained based on the above structure is high, and the generalization ability is strong.
[0029] A dual-branch structure of a feature extraction branch and a feature generation branch is adopted. Among them, the feature extraction branch is used to extract multi-scale preliminary features from the obtained multi-source remote sensing image data. The feature extraction branch selects the lightweight convolutional neural network MobileNet V2 as the backbone network and removes the final pooling layer and fully connected layer, which is used to extract multi-scale preliminary features from the remote sensing image data to reduce the model calculation complexity while ensuring the feature extraction ability, and improve the efficiency and deployment flexibility of the detection system. The MobileNet V2 network has the advantages of compact structure, small number of parameters, and fast calculation speed, and is especially suitable for deployment and use on devices with limited computing resources.
[0030] The feature generation branch is composed of multiple dense generation modules, which are used to perform fine-grained texture restoration and edge information enhancement on the initial image features, obtain multi-scale fine-grained features, realize the fine-grained texture restoration and edge information enhancement of the multi-scale preliminary features, and improve the detection ability of small targets and boundary structures. The feature generation branch can reconstruct fine-grained detail information from low-quality remote sensing images, make up for the information loss caused by low resolution in the feature extraction process, and can also enhance the texture features of the target area, making the detection results more accurate. A large number of skip connections in the multiple dense generation modules enhance the feature transmission and reuse ability, thereby strengthening the generation ability of detail features and reducing the impact caused by the low resolution of remote sensing images.
[0031] Adopting a dual-branch structure can make full use of the advantages of different branches, improve the classification and recognition ability of the model, enhance the robustness of the model, enable it to better process various input data, and the multi-branch structure can extract multiple features from the same input, increase the diversity of features, alleviate the overfitting problem of the model, and can also adapt to different task requirements.
[0032] In order to effectively supplement the fine-grained features extracted by the feature generation branch into the preliminary features obtained by the feature extraction branch, improve the integrity and discriminative ability of feature representation, a cross-attention fusion module is designed. Through the synergistic effect of spatial attention and channel attention, the features of the two branches are weighted and cross-fused to enhance the feature expression ability, strengthen the correlation and complementarity between the two types of features, and achieve more accurate feature fusion. In order to make full use of the obtained multi-scale features during the decoding process and make full use of multi-scale information, the present invention also designs a feature adaptive fusion module to gradually fuse high-level features into low-level features during the decoding process to improve the detection accuracy. Therefore, the feature fusion mechanism of the present invention includes a cross-attention fusion module and a feature adaptive fusion module.
[0033] The structure of the cross-attention fusion module is as Figure 3 shown, which is used to supplement the multi-scale fine-grained feature information into the multi-scale preliminary features, enhance the semantic expression ability of the features, and the cross-attention fusion mechanism can dynamically focus on the correlation between different modalities. By highlighting key features and suppressing redundant information through the cross-attention mechanism, the effective fusion of the semantic information of the feature extraction branch and the fine-grained features of the feature generation branch is realized, thereby enhancing the feature representation ability, realizing the efficient utilization of the dual-branch features, and improving the detection accuracy.
[0034] The structure of the feature adaptive module is as Figure 4As shown, during the decoding process, through dilated convolution and an adaptive guidance mechanism, high-level features are aggregated towards low-level features to achieve the complementary fusion of detailed information and semantic information, making full use of multi-level features, thereby improving the detection performance and generating a feature map with rich semantic expression capabilities. In the remote sensing image detection model, a ladder structure is adopted to stack the feature adaptive modules layer by layer, thereby realizing the fusion of adjacent-scale features and the feature interaction of non-adjacent scales, comprehensively capturing the detailed information of mining area targets, ensuring that the final feature map can retain complete and accurate semantic information, and thus improving the detection accuracy.
[0035] Through the fusion of cross-attention in feature adaptive fusion, features of different modalities are weighted and fused, and the weights are dynamically determined by the cross-attention mechanism, reflecting the importance of different features; combining the cross-attention fusion mechanism and the feature adaptive fusion mechanism can enhance the feature expression ability, improve the accuracy and robustness of the model, and can be applied to a variety of multi-modal tasks.
[0036] In one embodiment, the MobileNet V2 network includes an input layer, an initial convolutional layer, multiple inverted residual structure modules, a global average pooling layer, and a fully connected layer; among them, the input layer is used to receive the preprocessed temporal remote sensing image data; the initial convolutional layer uses a convolutional kernel of size 3×3 and a stride of 2 for initial feature extraction of the input image; the inverted residual structure module includes an expansion layer, a depthwise separable convolutional layer, and a projection layer, where the expansion layer is a 1×1 convolutional operation for expanding the channel dimension of the input features; the depthwise separable convolutional layer consists of a 3×3 depth convolution and a 1×1 pointwise convolution, where the depth convolution is used to independently perform spatial convolution operations on each channel, and the pointwise convolution is used for information integration between channels; the projection layer uses a 1×1 convolution to compress the high-dimensional features back to the original low-dimensional space and fuse them with the input features through skip connections; multiple inverted residual structure modules are stacked with different channel numbers and resolutions to construct a deep feature extraction network; the global average pooling layer is used to compress the feature map in the spatial dimension; the fully connected layer is used to map the global features to the classification results or detection outputs.
[0037] Specifically, the MobileNet V2 network includes an input layer, an initial convolutional layer, multiple inverted residual structure modules, a global average pooling layer, and a fully connected layer. The input layer is used to receive the preprocessed temporal remote sensing image data; the initial convolutional layer uses a convolutional kernel of size 3×3 and a stride of 2 for initial feature extraction of the input image and reducing the spatial resolution to support the computational efficiency of subsequent deep networks.
[0038] The inverted residual structure module is the core structure of MobileNet V2, mainly including an expansion layer, a depthwise separable convolutional layer, and a projection layer. Among them, the expansion layer is a 1×1 convolutional operation, which is used to expand the input features in the channel dimension and improve the feature expression ability; the depthwise separable convolutional layer consists of a 3×3 depth convolution and a 1×1 pointwise convolution. The former independently performs spatial convolution operations on each channel, and the latter is used for information integration between channels. This design significantly reduces the computational amount of the model and improves the efficiency of spatial feature extraction; the projection layer also uses a 1×1 convolution to compress the high-dimensional features back to the original low-dimensional space and fuse them with the input features through skip connections, thereby forming a complete inverted residual connection structure. Multiple inverted residual structure modules are stacked with different channel numbers and resolutions to construct a deep feature extraction network.
[0039] In addition, the feature extraction branch also includes a global average pooling layer, which is used to compress the high-dimensional feature map in the spatial dimension and extract global semantic information; finally, the global features are mapped to the task output space through a fully connected layer, and the final detection result is output.
[0040] By introducing MobileNet V2 as the backbone network of the feature extraction branch, it is possible to effectively control the volume and computational resource consumption of the model while maintaining a high detection accuracy, and it has good practicability and deployment adaptability.
[0041] In one embodiment, the dense generation module includes three residual dense blocks. Each residual dense block includes five 3×3 convolutional layers connected in series. After each convolutional operation, Leaky ReLU is used as the activation function, and there is no batch normalization layer; the convolutional layers are connected by skip connections through channel dimension concatenation, and the residual dense blocks are fused by element-wise addition, and a residual scaling coefficient is introduced to scale the fusion result; at the end of the structure of the dense generation module, a 3×3 convolutional layer with a stride of 2 is used for downsampling to make the size of the feature map consistent with the feature extraction branch; where the formula of the dense generation module is: ; In the formula, represents a convolutional operation with a kernel size of 3 and a stride of 1; represents the output result of the th convolutional operation in the current residual dense block; represents that before the i-th convolution in the current residual dense block, the features of the previous layers are concatenated in the channel dimension; respectively represent the output results of the three residual dense blocks; represents the Leaky ReLU activation function.
[0042] Specifically, as Figure 5As shown, each dense generation module consists of three sequentially stacked Residual Dense Blocks (RDBs). Each residual dense block contains five 3×3 convolutional layers connected in sequence. Skip connections are implemented through concatenation in the channel dimension between convolutional layers to fully achieve cross-layer feature transfer and reuse of low-level information. After each convolutional operation, the Leaky ReLU function is used as the activation function to enhance the non-linear expression ability. At the same time, batch normalization layers are not introduced in the entire structure to reduce redundant calculations during the training phase and improve the adaptability of the model in complex remote sensing scenarios.
[0043] Inside the residual dense block, the input of each convolutional operation is the concatenation result of the outputs of all previous layers, thus forming a typical dense connection structure. Feature fusion is achieved by element-wise addition between the three residual dense blocks to enhance the overall expression ability. To control the scale of the fused features, a residual scaling coefficient β is introduced after feature fusion to scale the fusion result, effectively suppressing gradient explosion and enhancing the stability of the training process.
[0044] In addition, to achieve spatial scale alignment with the feature extraction branch, a convolutional layer with a stride of 2 and a kernel size of 3×3 is set at the end of the dense generation module to downsample the feature map so that its size is consistent with the output of the feature extraction branch, facilitating subsequent multi-branch feature fusion operations.
[0045] Through the collaborative design of a large number of skip connections and dense residual structures, the dense generation module realizes the effective fusion of low-level texture information and high-level semantic information in remote sensing images, strengthens the ability to preserve edges and local structures, and thus improves the recognition ability and detection accuracy of the overall model for complex patch structures, and enhances the robustness of the model in different types of open-pit mining area scenarios.
[0046] In one embodiment, the cross-attention fusion module includes a spatial attention block and a channel attention block. Through the spatial attention block and the channel attention block, attention operations are performed on the multi-scale preliminary features and multi-scale fine-grained features in the spatial dimension and the channel dimension respectively to obtain a channel attention sequence and a spatial attention feature map, and a cross-matrix multiplication operation is performed to obtain two weight matrices. Based on the two weight matrices, element-wise multiplication re-weighting operations are performed on the multi-scale preliminary features and multi-scale fine-grained features respectively. An element-wise addition feature fusion operation is performed on the re-weighted features to obtain fused features. The fused features are adjusted through a depthwise separable convolutional layer to obtain the output features after fusing the features of the two branches.
[0047] Specifically, as Figure 3As shown in the figure, the cross-attention fusion module includes a spatial attention block (SAB) and a channel attention block (CAB). Through the collaborative action of spatial attention and channel attention, the correlation and complementarity between the two types of features are strengthened to achieve more accurate feature fusion.
[0048] By performing attention operations on the spatial dimension and channel dimension of the multi-scale preliminary features output by the feature extraction branch and the multi-scale fine-grained features output by the feature generation branch respectively, then performing a cross-matrix multiplication operation on the obtained channel attention sequence and spatial attention feature map to obtain two weight matrices, and based on the two weight matrices, performing element-wise multiplication on the multi-scale preliminary features and multi-scale fine-grained features respectively to re-weight them. Subsequently, an element-wise addition operation is performed on the weighted features to achieve feature fusion. Finally, the fused features are adjusted through a depthwise separable convolutional layer to obtain the final output features after fusing the features of the two branches.
[0049] Among them, the channel attention block is used to perform global max pooling and global average pooling on the input feature map in the spatial dimension respectively to obtain two groups of different channel information sequences; the two groups of channel information sequences are input into a multi-layer perceptron (MLP) through the channel attention block for feature reconstruction to extract high-level channel features; the two groups of channel features after feature reconstruction are fused through element-wise addition to obtain an output channel attention sequence containing rich channel information of the input feature map; among them, the formula of the channel attention block is: ; In the formula, and respectively represent the global average pooling and global max pooling operations, MLP represents the multi-layer perceptron, represents the input feature, represents the output sequence of the channel attention module.
[0050] Specifically, as Figure 4 shown, the channel attention block includes performing global max pooling (GAP) and global average pooling (GAP) on the input feature map in the spatial dimension respectively to obtain two groups of different channel information sequences; inputting the two groups of channel information sequences into a multi-layer perceptron for feature reconstruction to extract high-level channel features; fusing the two groups of channel features after reconstruction through element-wise addition to obtain an output channel attention sequence containing rich channel information of the input feature map.
[0051] Among them, the spatial attention block is used to enhance the semantic representation ability of the feature map. By performing global max pooling and global average pooling operations on the channel dimension of the input feature map respectively, and combining the global max pooling and global average pooling operations on the channel dimension of the input feature map, and compressing the channel dimension of the input feature map by means of 1×1 convolution, three spatial information matrices are generated; the three spatial information matrices are concatenated on the channel dimension and compressed through a convolution operation to obtain the output spatial information matrix , which is used to guide the feature enhancement of important spatial regions, and the formula is: ; ; ; ; In the formula, and respectively represent the max pooling operation and the average pooling operation on the channel dimension; represents the output spatial information matrix of the channel attention module; represents the point convolution operation with a convolution kernel size of 1; represents the concatenation operation on the channel dimension.
[0052] Specifically, the spatial attention block in the cross-attention fusion module is used to enhance the semantic representation ability of the feature map. Its processing process includes: First, perform global max pooling and global average pooling operations on the channel dimension of the input feature map respectively, and compress the channel dimension of the input feature map by means of 1×1 convolution to generate three spatial information matrices; then, concatenate the three spatial information matrices on the channel dimension and compress them through another convolution operation to obtain the output spatial information matrix, which is used to guide the feature enhancement of important spatial regions, thereby improving the expression ability of the feature fusion result.
[0053] In one embodiment, the feature adaptive fusion module is used to receive two sets of feature maps of adjacent scales, namely the low-level feature map and the high-level feature map ; perform bilinear interpolation upsampling processing on the high-level feature map through the feature adaptive module to make its spatial resolution consistent with that of the low-level feature map; perform 3×3 convolution on the upsampled high-level feature map to make its channel number consistent with that of the low-level feature map, and perform element-wise addition operation with the low-level feature map to perform preliminary feature fusion, and the formula is: ; In the formula, represents the bilinear interpolation upsampling operation; for the fused feature Three 3×3 dilated convolution operations with dilation rates of 1, 3, and 5 are respectively adopted to obtain feature representations under three different receptive fields, and they are concatenated in the channel dimension. Then, a 3×3 convolution is used to restore the channel dimension. The process is as follows: ; ; In the formula, represents the convolution operation with a dilation rate of . For the low-level feature map , feature compression is performed through a 1×1 convolution combined with the Sigmoid activation function to obtain the pixel-level guidance mask , and its expression is: ; By subtracting the matrix with all element values of 1 from , the anti-mask is obtained. Among them, and are respectively used to represent the context information of the target region and the background region. After concatenating and in the channel dimension, they are input into a 1×1 convolution and the Sigmoid activation function to generate a pixel weight matrix with the same resolution and number of channels as , which is used to recalibrate the features of to obtain the self-guided corrected features . The calculation formula is: ; In the formula, represents taking the inverse of the mask , that is, subtracting from the matrix . The output of the feature adaptive fusion module is obtained by weighted fusion of the aligned and corrected high-level features and the self-guided corrected low-level features , and then feature integration is performed through a 3×3 convolution. The formula is: ; In the formula, is the output of the feature adaptive fusion module.
[0054] Specifically, the feature adaptive fusion module receives the low-level feature map and the high-level feature map at adjacent scales, performs bilinear interpolation upsampling on the high-level feature map to make its spatial resolution consistent with that of the low-level feature map; performs 3×3 convolution on the upsampled high-level feature map to adjust the number of channels to be the same as that of the low-level feature map, and performs element-wise addition operation with the low-level feature map to achieve preliminary feature fusion. In order to reduce the spatial position information error that may occur during the feature fusion process, 3×3 dilated convolution operations with dilation rates of 1, 3, and 5 are respectively adopted for the fused features to obtain feature representations under three different receptive fields, and they are concatenated in the channel dimension, and then 1×1 convolution is used to restore the channel dimension.
[0055] For the low-level feature map , first, feature compression is performed through a 1×1 convolution combined with the Sigmoid activation function to obtain a pixel-level guidance mask. By subtracting the matrix (all its element values are 1) from , its inverse mask is obtained; after concatenating the mask and the inverse mask in the channel dimension, they are input into a 1×1 convolution and the Sigmoid activation function to generate a pixel weight matrix with the same resolution and number of channels as , which is used to recalibrate the features of to obtain self-guided corrected features; finally, the output of the feature adaptive fusion module is obtained by weighted fusion of the aligned and corrected high-level features and the self-guided corrected low-level features , and then feature integration is performed through 3×3 convolution.
[0056] Through the above steps, the complementary enhancement between features at different scales is effectively realized, and the detection performance of the target area in the remote sensing image is improved.
[0057] In one embodiment, after step S130 and before step S140, it further includes: performing weighted fusion on the binary cross-entropy loss and the Dice loss to construct a joint loss function, and the formula is: ; ; ; In the formula, and respectively represent the maximum values of and , and respectively represent the predicted value and the true label at the position in the image; the function represents the cross - entropy loss function, the symbol represents norm; the feature adaptive fusion modules are stacked step - by - step to produce three output results with the same resolution as the lowest - level feature map; the loss values of the three output results are calculated through the joint loss function, and after different weight coefficients are assigned to the three calculated loss values and weighted summation is performed, the final loss value is obtained, and the formula is: ; ; In the formula, , and respectively represent the loss values calculated from the output results of the three lowest - level feature adaptive fusion modules, , and are respectively the weight coefficients assigned to the corresponding loss terms, is the final loss value.
[0058] Specifically, in order to effectively evaluate the model performance and improve the detection accuracy, during the model training stage, the change prediction results output by the model are compared with the ground - truth labels, the loss value of each output is calculated, and the model parameters are optimized accordingly. Since the open - pit mining area usually accounts for a small proportion in the remote - sensing image, there is an obvious class - imbalance problem. If only the binary cross - entropy loss function is used, it is easy to cause the loss function to oscillate and the training process to be unstable. Therefore, the Dice loss function is introduced in the present invention to enhance the sensitivity and robustness of the model to small - target regions.
[0059] Furthermore, in order to balance the classification accuracy and the small - target detection ability, the binary cross - entropy loss and the Dice loss are weighted and fused to construct a joint loss function. Considering that the model structure contains multiple feature adaptive fusion modules, during the training process, the independent losses of the feature maps output by each module are calculated respectively, and then weighted summation is performed according to the set weight coefficients to form the overall loss function, so as to realize the in - depth supervision and collaborative optimization of the feature learning at each stage of the model.
[0060] Since the feature adaptive fusion modules are stacked step - by - step, three output results with the same resolution as the lowest - level feature map will be generated during the forward propagation process of the model. To further improve the convergence speed and stability during the model training process, the corresponding loss values of the above three outputs are calculated independently through the joint loss function. The final total loss function is obtained by weighted summation after different weight coefficients are assigned to the three loss values, so as to obtain a comprehensive optimization objective for further optimizing the model according to the loss function.
[0061] Step S140, inputting the remote sensing image data to be tested into the remote sensing image detection model to obtain the detection result of the open-pit mining area.
[0062] Specifically, after the remote sensing image detection model training is completed, the remote sensing image data to be tested is input into the remote sensing image detection model to generate a binary result map, and the remote sensing open-pit mine detection results are obtained based on the binary result map. The detection results are analyzed and interpreted, and the mining area in the remote sensing image is identified, thereby obtaining information on resource exploration, mineral management, and environmental monitoring.
[0063] Through the trained remote sensing image detection model, the remote sensing image data to be tested is inferred, and the output results after multi-scale fusion are combined to obtain a refined mining area detection map. Then, based on the mining area detection results in all remote sensing images, the spatial distribution information of the open-pit mine area can be extracted and the structural characteristics analyzed.
[0064] In this embodiment, a multi-phase remote sensing image dataset is constructed by acquiring multi-source high-resolution remote sensing image data at multiple time points covering the entire range of the target open-pit mining area, and preprocessing such as atmospheric correction, radiation calibration, color correction, image registration and cropping operations is performed to obtain time series remote sensing image data to improve image quality and time series consistency; an open-pit mine area detection model based on a dual-branch structure and a feature fusion mechanism is constructed by a deep learning algorithm, and is trained with time series remote sensing image data to obtain a remote sensing image detection model, wherein the dual-branch structure comprises a feature extraction branch and a feature generation branch, and the feature fusion mechanism comprises a cross-attention fusion module and a feature adaptive fusion model, thereby effectively realizing dynamic aggregation and expression enhancement of features of different scales and different semantic levels, and inputting the remote sensing image data to be tested into the remote sensing image detection model, and the model passes The mining area information in the remote sensing image is modeled from the two perspectives of spatial details and semantic structure through a dual-branch structure, and two feature fusion modules are used to realize information interaction and supplement between branches, so as to further improve the detection accuracy and model generalization ability, and obtain the detection results of the open-pit mining area. Combining the dual-branch structure and the feature fusion mechanism, it is possible to deeply mine the structural information and detail features in the remote sensing image, strengthen the feature fusion expression, improve the model's recognition ability for complex objects, ensure the recognition accuracy of small target details and mining area boundaries in low-resolution images, thereby improving the accuracy and robustness of remote sensing image detection, and can be applicable to target detection tasks in multi-scale and diversified remote sensing scenarios. In addition, the present invention has a low parameter amount and can achieve efficient calculation, so as to take into account the calculation efficiency while ensuring the detection accuracy, and improve the deployment efficiency and practicality of the model.
[0065] like Figure 2As shown in the figure, it is a schematic diagram of the detection effect of the algorithm of the present invention and existing algorithms such as GT (Gradient Boosting Tree), U-net, SegNet, DU SegNet, GIPNet, ASNet, Easy-Net, and DASFNet on targets with building interference, abnormal exposure, small targets, and complex edge structures. Obviously, the present invention can perform high-precision detection on targets in an interference environment with abnormal exposure, and is also applicable to the detection of tiny targets and targets with complex edge structures.
[0066] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0067] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a computer storage medium (ROM / RAM, magnetic disk, optical disk) and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. Therefore, the present invention is not limited to any specific combination of hardware and software.
[0068] The above content is a further detailed description of the present invention in combination with specific implementation manners. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should all be regarded as belonging to the protection scope of the present invention.
Claims
1. A remote sensing open-pit mining area detection method based on a dual-branch structure and a feature fusion mechanism, characterized in that Including the following steps: Obtain multi-source high-resolution remote sensing image data of the open-pit mine mining area; Preprocess the remote sensing image data to obtain time-series remote sensing image data. The preprocessing includes atmospheric correction, radiometric calibration, color correction, image registration, and cropping operations; Construct an open-pit mining area detection model based on a dual-branch structure and a feature fusion mechanism using a deep learning algorithm, and train it with the time-series remote sensing image data to obtain a remote sensing image detection model. The dual-branch structure includes a feature extraction branch and a feature generation branch, and the feature fusion mechanism includes a cross-attention fusion module and a feature adaptive fusion module; Among them, the feature extraction branch uses the lightweight MobileNet V2 as the backbone network and removes the final pooling layer and fully connected layer to extract multi-scale preliminary features from the remote sensing image data; The feature generation branch is composed of multiple dense generation modules, which are used to generate multi-scale fine-grained features from the remote sensing image data to enhance the learning of the mining area edge and texture details; The cross-attention fusion module is used to supplement the multi-scale fine-grained features to the multi-scale preliminary features; The feature adaptive fusion module is stacked in a ladder structure to perform fusion of adjacent-scale features and feature interaction of non-adjacent scales during the decoding process to comprehensively capture the detailed information of the mining area target. Input the to-be-detected remote sensing image data into the remote sensing image detection model to obtain the detection result of the open-pit mine mining area.
2. The remote sensing open-pit mining area detection method based on a dual-branch structure and a feature fusion mechanism according to claim 1, wherein The MobileNet V2 network includes an input layer, an initial convolutional layer, multiple inverted residual structure modules, a global average pooling layer, and a fully connected layer; Among them, the input layer is used to receive the preprocessed time-series remote sensing image data; The initial convolutional layer uses a convolutional kernel with a size of 3×3 and a stride of 2 to perform preliminary feature extraction on the input image; The inverted residual structure module includes an expansion layer, a depthwise separable convolutional layer, and a projection layer. Among them, the expansion layer is a 1×1 convolutional operation used to expand the channel dimension of the input features; the depthwise separable convolutional layer consists of a 3×3 depth convolution and a 1×1 pointwise convolution. The depth convolution is used to perform spatial convolution operations on each channel independently, and the pointwise convolution is used to integrate information between channels; the projection layer uses a 1×1 convolution to compress the high-dimensional features back to the original low-dimensional space and fuse them with the input features through a skip connection; Multiple inverted residual structure modules are stacked with different channel numbers and resolutions to construct a deep feature extraction network; The global average pooling layer is used to compress the feature map in the spatial dimension; The fully connected layer is used to map the global features to the classification result or detection output.
3. The remote sensing open-pit mining area detection method based on a dual-branch structure and a feature fusion mechanism according to claim 1, wherein The dense generation module includes three residual dense blocks, and each residual dense block includes five sequentially connected 3×3 convolutional layers. After each convolutional operation, Leaky ReLU is used as the activation function, and there is no batch normalization layer; Skip connections are made between the convolutional layers through concatenation in the channel dimension. Feature fusion is performed between the residual dense blocks by element-wise addition, and a residual scaling coefficient is introduced to scale the fusion result. At the end of the structure of the dense generation module, a 3×3 convolutional layer with a stride of 2 is used for downsampling to make the size of the feature map consistent with that of the feature extraction branch. Among them, the formula of the dense generation module is: ; In the formula, represents a convolution operation with a convolution kernel size of 3 and a stride of 1; represents the output result of the -th convolution operation in the current residual dense block; represents that in the current residual dense block, before the i-th convolution, the features of each previous layer are concatenated in the channel dimension; respectively represent the output results of three residual dense blocks; represents the Leaky ReLU activation function.
4. The remote sensing open-pit mining area detection method based on a dual-branch structure and a feature fusion mechanism according to claim 2, characterized in that, The cross-attention fusion module includes a spatial attention block and a channel attention block. Through the spatial attention block and the channel attention block, attention operations are respectively performed on the multi-scale preliminary features and the multi-scale fine-grained features in the spatial dimension and the channel dimension to obtain a channel attention sequence and a spatial attention feature map, and a cross-matrix multiplication operation is performed to obtain two weight matrices. Based on the two weight matrices, reweighting operations of element-wise multiplication are respectively performed on the multi-scale preliminary features and the multi-scale fine-grained features. Feature fusion operations of element-wise addition are performed on the reweighted features to obtain fused features. The fused features are adjusted through the depthwise separable convolutional layer to obtain the output features after fusing the features of the two branches.
5. The remote sensing open-pit mining area detection method based on a dual-branch structure and a feature fusion mechanism according to claim 4, characterized in that, The channel attention block is used to perform global max pooling and global average pooling respectively on the input feature map in the spatial dimension to obtain two groups of different channel information sequences. Through the channel attention block, the two groups of channel information sequences are input into a multi-layer perceptron for feature reconstruction to extract high-level channel features. The two groups of channel features after feature reconstruction are fused through element-wise addition to obtain an output channel attention sequence containing rich channel information of the input feature map. Among them, the formula of the channel attention block is: ; In the formula, and represent global average pooling and global max pooling operations respectively, MLP represents a multi-layer perceptron, represents the input feature, represents the output sequence of the channel attention module.
6. The remote sensing open-pit mining area detection method based on a dual-branch structure and a feature fusion mechanism according to claim 4, wherein The spatial attention block is used to enhance the semantic representation ability of the feature map. Through the spatial attention block, global max pooling and global average pooling operations are respectively performed on the channel dimension of the input feature map, and a 1×1 convolution is combined to compress the channel dimension of the input feature map to generate three spatial information matrices. Concatenate the three spatial information matrices in the channel dimension and compress them through a convolution operation to obtain the output spatial information matrix , which is used to guide the feature enhancement of important spatial regions. The formula is as follows: ; ; ; ; In the formula, and respectively represent the max pooling operation and the average pooling operation in the channel dimension; represents the output spatial information matrix of the channel attention module; represents the point convolution operation with a convolution kernel size of 1; represents the concatenation operation in the channel dimension.
7. The remote sensing open-pit mining area detection method based on a dual-branch structure and a feature fusion mechanism according to claim 1, characterized in that The feature adaptive fusion module is used to receive two sets of feature maps at adjacent scales, namely the low-level feature map and the high-level feature map ; The high-level feature map is upsampled by bilinear interpolation through the feature adaptive module to make its spatial resolution consistent with that of the low-level feature map. Perform a 3×3 convolution on the upsampled high-level feature map to make its number of channels consistent with that of the low-level feature map, and perform an element-wise addition operation with the low-level feature map for preliminary feature fusion. The formula is as follows: ; In the formula, represents the bilinear interpolation upsampling operation; For the fused features 3×3 dilation convolution operations with dilation rates of 1, 3, and 5 are respectively used to obtain feature representations under three different receptive fields, and they are concatenated in the channel dimension, and then a 3×3 convolution is used to restore the channel dimension. The process is as follows: ; ; In the formula, represents a convolution operation with a dilation rate of ; For the low-level feature map , feature compression is performed by a 1×1 convolution combined with the Sigmoid activation function to obtain a pixel-level guidance mask , and its expression is: ; By subtracting the matrix with all element values equal to 1 from , the anti-mask is obtained, where and are respectively used to represent the context information of the target region and the background region; After splicing and in the channel dimension, input them into the 1×1 convolution and the Sigmoid activation function to generate a set of pixel weight matrices with the same resolution and number of channels as and which are used to recalibrate the features of to obtain self-guided corrected features . The calculation formula is: a pixel weight matrix with the same resolution and number of channels for feature recalibration to obtain self-guided corrected features The calculation formula is: ; In the formula, represents taking the inverse of the mask , that is, subtracting from the matrix ; The output of the feature adaptive fusion module is obtained by weighted fusion of the high-level features after alignment correction and the low-level features after self-guided correction and is obtained after feature integration through 3×3 convolution. The formula is: ; In the formula, is the output of the feature adaptive fusion module.
8. The remote sensing open-pit mining area detection method based on a dual-branch structure and a feature fusion mechanism according to claim 1, wherein After constructing an open-pit mining area detection model based on a dual-branch structure and a feature fusion mechanism using a deep learning algorithm and training it with the time-series remote sensing image data to obtain a remote sensing image detection model, before inputting the remote sensing image data to be measured into the remote sensing image detection model to obtain the detection result of the open-pit mining area, it also includes: The binary cross-entropy loss and the Dice loss are weighted and fused to construct a joint loss function, and the formula is as follows: ; ; ; In the formula, and respectively represent and the maximum values, and respectively represent the predicted value and the true label at the position in the image; the function represents the cross-entropy loss function, and the symbol represents norm; The feature adaptive fusion module is stacked step by step to generate three output results with the same resolution as the lowest-level feature map. The loss values of the three output results are calculated through the joint loss function, and the calculated three loss values are weighted and summed after being assigned different weight coefficients to obtain the final loss value. The formula is: ; ; In the formula, , and respectively represent the loss values calculated from the output results of the three lowest-level feature adaptive fusion modules, , and are respectively the weight coefficients assigned to the corresponding loss terms, is the final loss value.
Citation Information
Patent Citations
Remote sensing image change detection method based on twinborn multi-scale cross attention
CN117975267A
Remote sensing mining pattern spot time sequence change detection method based on double-branch depth supervision
CN118230159A
Three-dimensional point-cloud semantic segmentation method based on multi-level boundary enhancement for unstructured environment
WO2024230038A1
Land change detection method combining matrix decomposition and adaptive propagation, and system
WO2025111921A1
Cited By
Landsat-MODIS remote sensing image space-time fusion method based on double-layer cross attention mechanism
CN121010871A
Ore image recognition method based on neural network
CN121033030A
Cross-scale space-time fusion ground feature classification method based on double-branch architecture
CN121121529A
A cross-scale spatio-temporal fusion feature classification method based on a double-branch architecture
CN121121529B
Information generation-oriented remote sensing image key infrastructure detection method and system
CN121170589A