Mine remote sensing scene classification method and device, electronic equipment and storage medium

By combining multiple remote sensing data modes and using feature enhancement and interactive aggregation network technology, the problem of low classification accuracy in mine remote sensing scenarios caused by relying on single mode data in the existing technology is solved, and higher classification accuracy and robustness are achieved.

CN120070941APending Publication Date: 2025-05-30新疆维吾尔自治区煤田地质局综合地质勘查队 +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411970990.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing mine remote sensing scenario classification methods rely on single-modal remote sensing data, resulting in low classification accuracy, especially in areas with complex terrain or dense vegetation.

Method used

Using a method combining the first image data and the second image data, the boundary features are enhanced by the first feature enhancement network and the second feature enhancement network, modal features are extracted through the feature extraction network, and fused through the feature interaction aggregation network, and finally scene classification is performed through a classifier.

Benefits of technology

By combining multiple remote sensing data modes, the information loss and random error caused by single mode data are reduced, and the accuracy and robustness of mine remote sensing scenario classification are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070941A_ABST
    Figure CN120070941A_ABST
Patent Text Reader

Abstract

The invention provides a mine remote sensing scene classification method and device, electronic equipment and a storage medium, and relates to the technical field of image processing, and the method comprises the steps: obtaining first image data and second image data corresponding to a target mountain area; enhancing the boundary feature of the first image data through a first feature enhancement network to obtain a first modal feature, and enhancing the boundary feature of the second image data through a second feature enhancement network to obtain a second modal feature; extracting the first modal feature through a first feature extraction network to obtain a first extraction feature, and extracting the second modal feature through a second feature extraction network to obtain a second extraction feature; fusing the first extraction feature and the second extraction feature through a feature interaction aggregation network to obtain a fused feature; and processing the fusion features through a classifier to obtain a scene classification result of the target mountain area. According to the method, various modal features can be combined, and the classification precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular, to a method, device, electronic device, and storage medium for classifying mine remote sensing scenes. Background Art

[0002] Traditional mine monitoring methods rely on ground surveys, manual inspections, etc. These methods are not only time-consuming and laborious, but also difficult to cover large areas. Especially under complex terrain conditions, the efficiency and accuracy are greatly limited. With the development of remote sensing technology, it is possible to classify mine scenes through modal remote sensing images obtained by satellites or drones. This can not only break through the limitations of the geographical environment and cover a wider area, but also improve the efficiency and accuracy of monitoring work through digital and automated means.

[0003] However, most of the mine remote sensing scene classification methods in the related art mainly rely on single-modal remote sensing data. Single-modal data can only provide single-dimensional features, and there may be problems such as random errors and sample biases, which may lead to poor accuracy in classifying mine remote sensing scenes. Especially in areas with complex terrain or dense vegetation, the data provided by single-modal data is limited, resulting in a reduction in mine classification accuracy. Summary of the Invention

[0004] The problem solved by the present invention is how to improve the classification accuracy of remote sensing scenes.

[0005] To solve the above problems, the present invention provides a method, device, electronic device, and storage medium for classifying mine remote sensing scenes.

[0006] In a first aspect, the present invention provides a method for classifying mine remote sensing scenes, including:

[0007] Obtaining first image data and second image data corresponding to a target mountain area;

[0008] Enhancing the boundary features of the first image data through a first feature enhancement network to obtain first modal features, and enhancing the boundary features of the second image data through a second feature enhancement network to obtain second modal features;

[0009] Extracting the first modal features through a first feature extraction network to obtain first extraction features, and extracting the second modal features through a second feature extraction network to obtain second extraction features;

[0010] Fusing the first extraction features and the second extraction features through a feature interaction aggregation network to obtain fusion features;

[0011] Processing the fusion features through a classifier to obtain a scene classification result of the target mountain area.

[0012] Optionally, enhancing the boundary features of the first image data through the first feature enhancement network to obtain first modal features, and enhancing the boundary features of the second image data through the second feature enhancement network to obtain second modal features, includes:

[0013] Processing the boundary features of the first image data through a first asymmetric convolution network to obtain first boundary enhancement features, and processing the boundary features of the second image data through a second asymmetric convolution network to obtain second boundary enhancement features;

[0014] Processing the first boundary enhancement features through a first patch-aware attention mechanism to obtain the first modal features, and processing the second boundary enhancement features through a second patch-aware attention mechanism to obtain the second modal features, where the first feature enhancement network includes the first asymmetric convolution network and the first patch-aware attention mechanism, and the second feature enhancement network includes the second asymmetric convolution network and the second patch-aware attention mechanism.

[0015] Optionally, processing the first boundary enhancement features through a first patch-aware attention mechanism to obtain the first modal features, and processing the second boundary enhancement features through a second patch-aware attention mechanism to obtain the second modal features, includes:

[0016] Extracting the first boundary enhancement features through a multi-branch feature extraction network to obtain first multi-branch extraction features, and extracting the second boundary enhancement features through the multi-branch feature extraction network to obtain second multi-branch extraction features, where both the first patch-aware attention mechanism and the second patch-aware attention mechanism include the multi-branch feature extraction network and an attention mechanism, and the multi-branch feature extraction network processes the first boundary enhancement features and the second boundary enhancement features in parallel through a local branch, a global branch, and a serial convolution branch, and adjusts the local branch and the global branch through a patch size parameter, and the patch size parameter represents the spatial size of the patch;

[0017] Performing feature enhancement processing on the first multi-branch extraction features through the attention mechanism to obtain the first modal features, and performing feature enhancement processing on the second multi-branch extraction features through the attention mechanism to obtain the second modal features.

[0018] Optionally, before extracting the first modal features through the first feature extraction network to obtain first extraction features, and extracting the second modal features through the second feature extraction network to obtain second extraction features, further includes:

[0019] Stack the convolutional layer and the attention layer to form a preset stacked network structure, which are respectively used as the first feature extraction network and the second feature extraction network, where the structures of the first feature extraction network and the second feature extraction network are the same.

[0020] Optionally, the preset stacked network structure includes at least two convolutional blocks and at least two self-attention mechanism blocks, where the convolutional block includes a first convolutional layer, a second convolutional layer, and a third convolutional layer. The first convolutional layer is used to expand features, the second convolutional layer is used to extract features, and the third convolutional layer is used to compress features; the self-attention mechanism block is connected to a feed-forward neural network, and the feed-forward neural network is used to expand and compress the features to be processed.

[0021] Optionally, the step of fusing the first extracted feature and the second extracted feature through the feature interaction aggregation network to obtain a fused feature includes:

[0022] Combine the first extracted feature and the second extracted feature to obtain a combined feature;

[0023] Perform feature correction on the combined feature through a correction weight matrix to obtain a corrected feature, where the correction weight matrix is used to adjust the channel importance weights of the combined feature;

[0024] Obtain the fused feature through the corrected feature.

[0025] Optionally, the step of obtaining the fused feature through the corrected feature includes:

[0026] Split the corrected feature into a first corrected feature and a second corrected feature in the channel dimension;

[0027] Obtain local attention weights and global attention weights through the corrected feature;

[0028] Obtain a first attention weight through the local attention weight and the global attention weight;

[0029] Add the first attention weight to the first corrected feature to obtain a first weighted feature;

[0030] Add a second attention weight to the second corrected feature to obtain a second weighted feature;

[0031] Add the first weighted feature and the second weighted feature to obtain the fused feature, where the sum of the first attention weight and the second attention weight is 1.

[0032] In a second aspect, the present invention provides a device for classifying mine remote sensing scenes, including:

[0033] An acquisition module, which is used to acquire first image data and second image data corresponding to a target mountain area;

[0034] A feature enhancement module, which is used to enhance the boundary features of the first image data through a first feature enhancement network to obtain first-modal features, and enhance the boundary features of the second image data through a second feature enhancement network to obtain second-modal features;

[0035] A feature extraction module, which is used to extract the first-modal features through a first feature extraction network to obtain first extraction features, and extract the second-modal features through a second feature extraction network to obtain second extraction features;

[0036] A feature interaction and aggregation module, which is used to fuse the first extraction features and the second extraction features through a feature interaction and aggregation network to obtain fused features;

[0037] A classification module, which is used to process the fused features through a classifier to obtain a scene classification result of the target mountain area.

[0038] In a third aspect, the present invention provides an electronic device, including a memory and a processor;

[0039] The memory is used to store a computer program;

[0040] The processor is used to implement the mine remote sensing scene classification method as described in the first aspect when executing the computer program.

[0041] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the mine remote sensing scene classification method as described in the first aspect is implemented.

[0042] The beneficial effects of the mine remote sensing scene classification method of the present invention are as follows: By combining two types of modal data, namely the first image data and the second image data, it is possible to reduce the limitations brought by single-modal data through different types of remote sensing information, improve the accuracy of the classification results, and help to more comprehensively understand the characteristics of the target area. By processing the first image data and the second image data respectively through the first feature enhancement network and the second feature enhancement network, the detection accuracy of the boundary features can be effectively improved. The boundary features are important spatial information in image analysis, and enhancing the boundary features helps to improve the effectiveness of subsequent feature extraction. After being processed by the first feature enhancement network and the second feature enhancement network, the first modal features and the second modal features obtained have class imbalance and noise, which will affect the classification effect. By extracting the first modal features and the second modal features respectively through the first feature extraction network and the second feature extraction network, important features can be extracted, the interference of noise can be reduced, and the influence brought by data imbalance and noise can be alleviated. By extracting the respective advantages of the first image data and the second image data and performing feature fusion through the feature interaction aggregation network, combining the first and second extracted features can effectively reduce the information loss caused by a single modality, avoid the problems of random errors and sample biases that may exist in single-modal data, and at the same time enhance the robustness of the model. By classifying through the classifier, it can meet the classification requirements of various remote sensing scenes, and provide more accurate classification results through comprehensive analysis of the fused features. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 It is a schematic flowchart of a mine remote sensing scene classification method according to an embodiment of the present invention;

[0044] Figure 2 It is a flowchart of a mine remote sensing scene classification method according to an embodiment of the present invention;

[0045] Figure 3 It is a detailed flowchart of a mine remote sensing scene classification method according to an embodiment of the present invention;

[0046] Figure 4 It is a schematic flowchart of the feature enhancement step of a mine remote sensing scene classification method according to an embodiment of the present invention;

[0047] Figure 5 It is a schematic flowchart of the feature interaction aggregation step of a mine remote sensing scene classification method according to an embodiment of the present invention;

[0048] Figure 6 It is an example diagram of a mine remote sensing scene classification device according to an embodiment of the present invention;

[0049] Figure 7 It is an example diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] To make the above objects, features, and advantages of the present invention more apparent and understandable, the following provides a detailed description of specific embodiments of the present invention with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments described herein. Instead, these embodiments are provided to more thoroughly and completely understand the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not used to limit the protection scope of the present invention.

[0051] It should be understood that the various steps described in the method embodiments of the present invention can be executed in different orders and / or executed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this regard.

[0052] As used herein, the term "including" and its variations are open-ended, that is, "including but not limited to"; the term "based on" is "at least partially based on"; the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments"; the term "optionally" means "optional embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts such as "first" and "second" mentioned in the present invention are only used to distinguish different devices, modules, or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules, or units.

[0053] It should be noted that the modifications of "one" and "multiple" mentioned in the present invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly stated in the context, it should be understood as "one or more".

[0054] The names of the messages or information exchanged between multiple devices in the embodiments of the present invention are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0055] This embodiment provides a method, apparatus, electronic device, and storage medium for classifying mine remote sensing scenes.

[0056] As Figure 1 and Figure 2 shown, a method for classifying mine remote sensing scenes provided by an embodiment of the present invention includes:

[0057] Step S100, obtaining first image data and second image data corresponding to a target mountain area.

[0058] In one embodiment, the first image data may be visual (Red, Green, Blue, RGB) image data, and the second image data may be Synthetic Aperture Radar (SAR) data. Visual images can provide rich color and texture information, clearly showing the color differences of surface cover types; synthetic aperture radar data has the ability to penetrate clouds and vegetation, enabling all-weather acquisition of ground object information and providing rich ground object structure information. Synthetic aperture radar data has all-weather monitoring capabilities and is not affected by weather conditions, but synthetic aperture radar data lacks color information and it is sometimes difficult to visually identify ground object categories. By combining with visual image data, the limitations of synthetic aperture radar data in this regard can be compensated, and at the same time, due to the strong penetration of synthetic aperture radar data, the influence of shadows and vegetation on the classification of mine scenes can be reduced. By combining the two types of data, more comprehensive ground object information of the target mountain area can be obtained.

[0059] Step S200: Enhance the boundary features of the first image data through a first feature enhancement network to obtain first-modal features, and enhance the boundary features of the second image data through a second feature enhancement network to obtain second-modal features.

[0060] Specifically, the first feature enhancement network and the second feature enhancement network may be Feature Enhancement Networks (FENet). The feature enhancement network extracts the information of the first image data and the second image data through a Feature Fusion Enhancement Module (FFEM) based on the correlation and complementarity of cross-modal features, enhances the recognition ability of significant regions through cross-fusion and hybrid channel attention mechanisms, and supplements shallow detail information and avoids interference from low-quality underlying information by introducing a Boundary Feature Enhancement Module (BFEM). Thus, it can more effectively capture the boundary features of the target.

[0061] The first feature enhancement network and the second feature enhancement network may enhance the boundary features of the first image data and the second image data through asymmetric convolution and patch-aware attention mechanisms. By using asymmetric convolution kernels to process the first image data and the second image data, boundary information in different directions can be effectively captured, improving the recognition accuracy of boundary features. At the same time, in combination with the patch-aware attention mechanism, multi-scale extraction of local and global features is achieved by adjusting the patch size parameter, and important features are enhanced through efficient channel attention and spatial attention modules, ensuring that the feature representation contains both rich boundary details and does not lose internal structure and context relationships.

[0062] Step S300: Extract the first-modal features through a first feature extraction network to obtain first-extracted features, and extract the second-modal features through a second feature extraction network to obtain second-extracted features.

[0063] Specifically, the first feature extraction network and the second feature extraction network can adopt an optimized CNN-Transformer network structure, which combines the local feature extraction ability of the convolutional neural network and the global context modeling ability of the Transformer, and can efficiently extract multi-scale features.

[0064] Step S400, through the feature interaction aggregation network, fuse the first extracted feature and the second extracted feature to obtain a fused feature.

[0065] Specifically, the features obtained from the first feature extraction network and the second feature extraction network are fused through the feature interaction aggregation network. The feature interaction aggregation network realizes the complementarity and enhancement of different modality features through the attention-based adaptive feature aggregation network, leveraging the advantages of the first image data and the second image data. Through the adaptively adjusted weight matrix, the model can dynamically adjust the attention to local and global features. Local features are related to fine-grained information in the image, such as texture, edges, corners, etc. In the presence of occlusion or deformation, local features can help the model focus on key parts, thereby improving the robustness of object recognition. Global features capture the structured information within the entire image or a larger area, such as shape, layout, and context relationships, which helps the model understand the overall structure of the scene and enhance the model's understanding ability of complex scenes.

[0066] In one implementation, the goal is to identify and classify specific types of ore bodies. Mines are often located in complex terrain environments with terrain features such as steep slopes and deep valleys. When conducting identification, adverse weather conditions such as foggy days and rainy days may be encountered, affecting the quality of remote sensing images. Vegetation cover, mechanical equipment, and temporary buildings may occlude the ore body or cause image errors. By combining the advantages of RGB image data (i.e., rich color and detailed expressiveness) and SAR image data (i.e., all-weather, all-day monitoring ability and strong penetration), and the feature interaction aggregation network, the RGB image features and SAR image features are fused and interacted to capture local features, such as fine-grained information like texture, edges, and corners. When some mining areas are blocked by trees, by enhancing the attention to the edge and texture features of the unblocked parts, the model can still identify the presence of the ore body and accurately locate its boundary. Global features are captured, that is, the structured information within the entire mine area or a larger range, such as shape, layout, and context relationships. When evaluating mine safety, the model not only needs to identify the location of the ore body but also understand the overall layout of the surrounding environment. For example, by analyzing global features, the distribution of roads, the location of buildings, and potential dangerous areas (such as high slopes) around the mine can be identified. Global features help the model understand the overall structure of the mine and evaluate the safety condition of the mine, such as detecting potential landslide risk areas.

[0067] Step S500, process the fused features through a classifier to obtain the scene classification result of the target mountain area.

[0068] Specifically, a classifier refers to a component in a machine learning or deep learning model used to map input data (such as images, text, etc.) to a predefined set of categories. By analyzing and processing the fused features through a classifier, the final classification result is obtained. The settings of the classifier can be adjusted according to the characteristics of the specific task and dataset. For example, different numbers of fully connected layers can be selected, different activation functions can be used, or a dropout layer can be added to prevent overfitting. In addition, techniques such as Label Smoothing can be introduced to improve the generalization ability of the classifier.

[0069] In this embodiment, by combining two types of modal data, namely the first image data and the second image data, the limitations brought by single-modal data can be reduced through different types of remote sensing information, the accuracy of the classification result can be improved, and it helps to more comprehensively understand the characteristics of the target area. By respectively processing the first image data and the second image data through the first feature enhancement network and the second feature enhancement network, the detection accuracy of the boundary features can be effectively improved. Boundary features are important spatial information in image analysis, and enhancing boundary features contributes to the effectiveness of subsequent feature extraction. After being processed by the first feature enhancement network and the second feature enhancement network, the first modal features and the second modal features obtained have class imbalance and noise, which will affect the classification effect. By respectively extracting the first modal features and the second modal features through the first feature extraction network and the second feature extraction network, important features can be extracted, the interference of noise can be reduced, and the impact brought by data imbalance and noise can be alleviated. By extracting the respective advantages of the first image data and the second image data and performing feature fusion through the feature interaction aggregation network, combining the first and second extracted features can effectively reduce the information loss caused by a single modality and at the same time enhance the robustness of the model. By classifying through a classifier, the classification requirements of various remote sensing scenarios can be adapted, and a more accurate classification result can be provided by comprehensively analyzing the fused features.

[0070] Optionally, the process of enhancing the boundary features of the first image data through the first feature enhancement network to obtain the first modal features and enhancing the boundary features of the second image data through the second feature enhancement network to obtain the second modal features includes:

[0071] Process the boundary features of the first image data through the first asymmetric convolution network to obtain the first boundary enhancement features, and process the boundary features of the second image data through the second asymmetric convolution network to obtain the second boundary enhancement features.

[0072] Process the first boundary-enhanced feature through a first patch-aware attention mechanism to obtain the first modal feature, and process the second boundary-enhanced feature through a second patch-aware attention mechanism to obtain the second modal feature, where the first feature enhancement network includes the first asymmetric convolution network and the first patch-aware attention mechanism, and the second feature enhancement network includes the second asymmetric convolution network and the second patch-aware attention mechanism.

[0073] Specifically, the first feature enhancement network performs feature enhancement through the first asymmetric convolution network and the first patch-aware attention mechanism. The first asymmetric convolution network captures the non-linear features of the boundaries in the first image data, enhances the clarity and contrast of the boundaries, and obtains the first boundary-enhanced feature. The asymmetric convolution network uses a multi-branch asymmetric convolution kernel. In one embodiment, as Figure 4 shown, use 1×3, 3×1, and 3×3 asymmetric convolution kernels to process the first image data to capture the multi-directional features of the boundaries. The 1×3 convolution can effectively detect horizontal line segments or structures close to the horizontal direction. F1 represents the feature obtained by capturing the edge information in the horizontal direction using the 1×3 convolution kernel; the 3×1 convolution kernel helps to identify vertical line segments or structures close to the vertical direction, such as building walls, road boundaries, etc. F2 represents the feature obtained by capturing the edge information in the vertical direction using the 3×1 convolution kernel; the 3×3 convolution kernel can enhance the recognition of oblique edges and other complex shapes. F3 represents the feature obtained by capturing the information (including the diagonal direction) in a wider local area of the diagonal using the 3×3 convolution kernel. Combine F1, F2, and F3 to obtain the first boundary-enhanced feature. Process the first boundary-enhanced feature processed by the first patch-aware attention mechanism and the asymmetric convolution network. Through the cooperation of the local branch, the global branch, and the serial convolution branch (convolutions of different patch sizes), realize the interaction and enhancement of local and global features. Through multi-branch feature extraction combined with channel and spatial attention for adaptive feature enhancement, obtain the first modal feature. Process the boundary feature of the second image data through the second asymmetric convolution network to obtain the second boundary-enhanced feature. Process the second boundary-enhanced feature through the second patch-aware attention mechanism to obtain the second modal feature.

[0074] In this optional embodiment, processing the first image data and the second image data through an asymmetric convolutional network can more accurately capture the boundary features of the mining area. The boundary enhancement features obtained after processing by the asymmetric convolutional network can accurately reflect the important information in the data, which helps subsequent feature extraction and classification. However, when the asymmetric convolutional network enhances the boundary features, it will ignore some directional boundary features. By adding a patch-aware attention mechanism and introducing multiple feature extraction paths (local, global, serial), the key regions of feature extraction can be adaptively adjusted to fully extract the important information in the features. By changing the size of the patch, local and global feature extraction are realized, making up for the shortcomings of the asymmetric convolutional network. Finally, channel and spatial attention are combined to perform adaptive feature enhancement on the extracted local and global features. The asymmetric convolutional network and the patch-aware attention mechanism achieve dual feature enhancement of the first image data and the second image data, ensuring that the features obtained from two different modalities can retain useful information to the greatest extent, reducing the influence of irrelevant information, and improving the overall performance and robustness of the model.

[0075] Optionally, processing the first boundary enhancement feature through the first patch-aware attention mechanism to obtain the first modality feature, and processing the second boundary enhancement feature through the second patch-aware attention mechanism to obtain the second modality feature includes:

[0076] Extracting the first boundary enhancement feature through a multi-branch feature extraction network to obtain a first multi-branch extraction feature, and extracting the second boundary enhancement feature through the multi-branch feature extraction network to obtain a second multi-branch extraction feature, where both the first patch-aware attention mechanism and the second patch-aware attention mechanism include the multi-branch feature extraction network and the attention mechanism. The multi-branch feature extraction network processes the first boundary enhancement feature and the second boundary enhancement feature in parallel through a local branch, a global branch, and a serial convolution branch, and adjusts the local branch and the global branch through a patch size parameter, and the patch size parameter represents the spatial size of the patch.

[0077] Performing feature enhancement processing on the first multi-branch extraction feature through the attention mechanism to obtain the first modality feature, and performing feature enhancement processing on the second multi-branch extraction feature through the attention mechanism to obtain the second modality feature.

[0078] Specifically, the multi-branch feature extraction network extracts features in parallel through different convolutional paths, including a local branch, a global branch, and a serial processing branch. The local branch is used to capture detailed features, enhancing the sensitivity to small features by focusing on the local details of the image; the global branch analyzes the information of the entire image and can integrate the overall information; the serial convolutional branch extracts features through patches of multiple scales, enhancing the ability to extract features of different scales. As Figure 4 shown, the local branch and the global branch are distinguished by controlling the patch size parameter p. When p = 2 is set, it means that the input features are divided into smaller patches of size 2x2. The local branch focuses on small regions of the image through a smaller patch size, ensuring that high-resolution local features such as edges and textures can be extracted. Since each patch covers a relatively small area, it can focus more on the structures and boundaries in the feature map. In contrast, when p = 4 is set, it means that larger patches of size 4x4 are used to process the features. The large-sized patches can cover a larger spatial range, thus capturing more extensive context information. The global branch uses a larger patch size to cover a wider area and capture global features such as the overall layout of the scene or large-scale structures. The serial convolutional branch gradually enhances the features through consecutive convolutional operations and can capture information at different levels. By adjusting the patch size, the spatial dimension aggregation and displacement of the patches are controlled, thereby achieving feature extraction at different scales. The patch size determines how the image is segmented in the spatial dimension and affects the depth of feature understanding of each branch. Smaller patches are more suitable for capturing fine-grained features, while larger patches help to understand the global context. The patch size parameter is achieved through the aggregation and displacement of non-overlapping patches in the spatial dimension. After feature extraction by the multi-branch feature extraction, through the attention matrix of the attention mechanism, different weights are assigned to achieve the extraction and interaction of local features and global features, fuse the multi-branch extracted features, and enhance the important features to obtain the first-mode features.

[0079] In this optional embodiment, by setting different patch size parameters in the multi-branch feature extraction network, effective capture of features at different scales is achieved, enabling more comprehensive capture of deep information in the image and ensuring that both subtle changes and overall structures can be effectively recognized. By extracting local features with small patches, detailed information and small target features in the boundaries can be effectively captured, enhancing the model's recognition ability for complex boundary regions. By extracting global information with large patches, the overall shape and semantic trends of the boundaries can be captured, enhancing the model's understanding ability of the global context. Through serial convolution operations, local and global features are integrated, effectively enhancing the correlation between local and global features and making the extracted features more discriminative and robust. Through the attention mechanism, the features extracted by multiple branches are weighted to further optimize the feature representation. Channel attention can highlight relevant feature channels while suppressing redundant or irrelevant channel features, thereby improving the expression efficiency of features. Combining channel and spatial attention for feature enhancement of the extracted local and global features.

[0080] Optionally, before extracting the first-modal features through the first feature extraction network to obtain first extraction features and extracting the second-modal features through the second feature extraction network to obtain second extraction features, it further includes:

[0081] Stacking a convolutional layer and an attention layer to form a preset stacked network structure, which are respectively used as the first feature extraction network and the second feature extraction network, where the structures of the first feature extraction network and the second feature extraction network are the same.

[0082] Specifically, the convolutional layer is an important part of the first feature extraction network and is used to construct a convolutional neural network (CNN). The convolutional neural network extracts local spatial information in the first-modal features through the local receptive field mechanism. The convolutional layer extracts different features of the image through the convolutional operations of multiple convolutional kernels. Stacking convolutional layers can form a deep network structure and enhance the feature extraction ability. The attention layer is constructed by the self-attention mechanism and is used to extract the global context information in the first-modal features. The attention layer processes features through the self-attention mechanism, which can dynamically adjust weights according to the importance of features, thereby improving the expression ability of features. The self-attention mechanism can effectively reduce the impact of noise on the classification effect. Stacking the constructed convolutional layer and attention layer in a preset manner forms a multi-layer feature extraction network. The stacking structure can adjust the number of layers and the specific configuration of each layer according to requirements, such as increasing or decreasing the number of convolutional layers and adjusting the position of the attention layer, etc., to optimize the network performance. The first feature extraction network not only has a strong ability to extract local features but also can effectively enhance the understanding of global features.

[0083] In this optional embodiment, the convolutional layer processes the first-modal features through the local receptive field mechanism. Through multiple convolutional operations, it captures detailed information such as edges and textures in the first-modal features and has strong recognition ability for small targets and complex textures. The self-attention mechanism can weight and integrate features in different spatial dimensions, realizing the effective fusion of local features and global features. The model can understand context information within a larger range. The self-attention mechanism enables the model to better capture the relationships between various parts of the image, can adaptively focus on important features, and suppress irrelevant features, thereby improving the robustness of the model and the accuracy of classification. By stacking the convolutional layer and the attention layer into a preset network structure, the effective combination of local feature extraction and global feature understanding is achieved. The convolutional layer in the stacked structure extracts local details, and the attention layer captures global semantic information. The model can simultaneously focus on detailed features and overall structural features. This not only increases the depth and expressive power of the network but also improves the robustness of the model and the accuracy of classification through a reasonable hierarchical architecture.

[0084] Optionally, the preset stacked network structure includes at least two convolutional blocks and at least two self-attention mechanism blocks. Among them, the convolutional block includes a first convolutional layer, a second convolutional layer, and a third convolutional layer. The first convolutional layer is used to expand features, the second convolutional layer is used to extract features, and the third convolutional layer is used to compress features; the self-attention mechanism block is connected to a feedforward neural network, and the feedforward neural network is used to expand and compress the features to be processed.

[0085] Specifically, each convolutional block includes three convolutional layers, namely the first convolutional layer, the second convolutional layer, and the third convolutional layer. The first convolutional layer uses a 1×1 convolutional kernel to expand the number of channels of the input features and enhance the expressive power of the input features; the second convolutional layer uses a 3×3 depth convolutional kernel to extract features, focuses on capturing local information in the image, performs local receptive field operations on the input features, and extracts information such as edges and textures in the features; the third convolutional layer uses a 1×1 convolution to compress features, compresses high-dimensional features into lower-dimensional features, reduces the number of channels to reduce computational complexity, and at the same time retains important features. The convolutional block uses a residual connection structure to avoid the problem of gradient disappearance and ensure the training stability of the deep network. The self-attention mechanism block globally models the features through the self-attention mechanism and is connected to a feedforward neural network (Feedforward Neural Network, FFN) to achieve the expansion and compression of features. The feedforward neural network consists of two fully connected layers. The first fully connected layer maps the input vector to a higher-dimensional feature space through a linear transformation, which helps to capture more complex feature information. The second fully connected layer then maps the high-dimensional feature space back to the initial dimension, thereby achieving the compression and integration of features.

[0086] In one implementation, the preset stacked network structure may consist of two convolutional blocks and two self-attention mechanism blocks. The first convolutional block includes a first convolutional layer, a second convolutional layer, and a third convolutional layer. Taking the first-modal features as the first input features, the first convolutional layer is used to expand the number of channels of the first input features, improve the expressive ability of the features, and obtain first expanded features; the second convolutional block is used to extract the local features of the first expanded features, capture details such as edges and textures of the image, and obtain first captured features; the third convolutional block is used to compress the first captured features to obtain first output features. Through residual connection, the first input features and the first compressed features are added together to obtain the first output features; the first output features are used as the input features of the second convolutional block, and the second compressed features are obtained through the processing of the second convolutional block. The first output features and the second compressed features are added together to obtain the second output features; the second output features are used as the input features of the first self-attention mechanism block, the second output features are expanded through the first fully connected layer to obtain third expanded features, the third expanded features are compressed through the second fully connected layer to obtain third compressed features, and through residual connection, the second output features and the third compressed features are added together to obtain third output features; the third output features are used as the input features of the second self-attention mechanism block, and fourth compressed features are obtained through the processing of the second self-attention mechanism block. The third output features and the fourth compressed features are added together to obtain first extracted features.

[0087] In this optional embodiment, through the collaborative action of the first convolutional layer, the second convolutional layer, and the third convolutional layer in the convolutional block, the expansion, extraction, and compression of the input features are realized. The first convolutional layer maps low-dimensional features to a high-dimensional space to enhance the representation ability of the features; the second convolutional layer extracts details such as textures and edges in the features through local convolution operations for feature extraction of small targets and complex boundaries in the data; the third convolutional layer reduces the computational complexity by compressing the feature dimension while enhancing the features. The self-attention mechanism block enhances the expressive ability of the features, reduces the interference of noise, enhances the model's ability to understand global context information, and significantly improves the accuracy and robustness of classification. The feed-forward neural network further enhances the expressive ability of the features and reduces feature redundancy by expanding and compressing the feature dimension. Stacking the convolutional block and the self-attention mechanism block realizes the efficient fusion of local features and global features. The local features extracted by the convolutional block can capture details such as edges and textures in the remote sensing image. Through the global modeling of the local features by the self-attention mechanism block, the model's ability to understand large-scale targets in the remote sensing scene is improved. The stacking of the convolutional block and the self-attention mechanism block enables the model to simultaneously focus on local details and global semantic features, improving the accuracy of classification.

[0088] Optionally, the fusion of the first extracted features and the second extracted features through the feature interaction aggregation network to obtain fused features includes:

[0089] Combine the first extracted feature and the second extracted feature to obtain a combined feature.

[0090] Perform feature calibration on the combined feature through a calibration weight matrix to obtain a calibrated feature, where the calibration weight matrix is used to adjust the channel importance weights of the combined feature.

[0091] Obtain the fused feature from the calibrated feature.

[0092] Specifically, as Figure 5 shown, by combining the first extracted feature and the second extracted feature, different modality features can be combined. For example, RGB image data provides rich texture and color information, and SAR image data has the ability to penetrate clouds and vegetation. The combination of the two can significantly improve the description ability for complex scenes. The first extracted feature and the second extracted feature are concatenated in the channel dimension to form a combined feature, thereby retaining the complete information of the two modality features. Perform feature calibration on the combined feature through a recalibration weight (recal_w) matrix to obtain a calibrated feature. The calibration weight matrix is used to adjust the importance weights of each channel of the combined feature, adaptively highlighting the key channel features and suppressing the redundant or noisy channel features.

[0093] In this optional embodiment, by combining features of different modalities, richer information can be obtained, improving the model's understanding ability for complex scenes. The calibration weight matrix dynamically adjusts the channel weights of the combined feature through an attention mechanism, which can adaptively highlight the important channel features and suppress the redundant or noisy channel features, thereby improving the efficiency and accuracy of feature expression. By calibrating the combined feature through the calibration weight matrix, the expression ability of the key channel features is significantly enhanced, enabling the model to pay more attention to the feature regions important for the classification task. The calibration operation can effectively suppress the redundant and noisy features in the combined feature, and the calibrated feature after calibration can better integrate the advantages of the two modality features, improving the discrimination ability for complex scenes, thereby improving the overall performance and robustness of remote sensing classification.

[0094] Optionally, the obtaining the fused feature from the calibrated feature includes:

[0095] Divide the calibrated feature into a first calibrated feature and a second calibrated feature in the channel dimension.

[0096] Obtain local attention weights and global attention weights from the calibrated feature.

[0097] Obtain a first attention weight from the local attention weights and the global attention weights.

[0098] Add the first attention weight to the first corrected feature to obtain a first weighted feature.

[0099] Add the second attention weight to the second corrected feature to obtain a second weighted feature.

[0100] Add the first weighted feature and the second weighted feature to obtain the fused feature, where the sum of the first attention weight and the second attention weight is 1.

[0101] Specifically, as Figure 5 shown, the corrected feature is split along the channel dimension into a first corrected feature and a second corrected feature. Global features (Fga) and local features (Fla) are generated from the corrected feature respectively. The global attention weight is generated based on the global feature, and the local attention weight is generated based on the local feature. The global attention weight extracts the global context information of the corrected feature through global pooling or self-attention mechanism to generate the attention weight in the global range, which is suitable for capturing the overall structural features of the scene. The local attention weight extracts the local region information of the corrected feature through a small convolutional kernel (such as 3×3) or local receptive field mechanism to generate the attention weight in the local region, which is suitable for capturing the features of small targets or boundary regions. The local attention weight and the global attention weight are fused, and through the Sigmoid activation function, the first attention weight (w) is generated. The first corrected feature is weighted by the first attention weight (w) to highlight the important features in the first corrected feature. The second corrected feature is weighted by the second attention weight (1 - w) to highlight the important features in the second corrected feature. The first weighted feature and the second weighted feature are added to generate the final fused feature, which contains the information of both local and global features.

[0102] In this optional embodiment, by splitting the corrected feature into two parts along the channel dimension, the model can process different features more carefully, enhancing the feature expression ability and avoiding interference between the two features, thus improving the expression ability of each feature. The calculation of local and global attention weights enables the model to dynamically adjust the focus according to the importance of the features, improving the effectiveness of the features so that the model can better capture key features and improve the classification accuracy. By generating local attention weights, it is possible to effectively capture detailed information such as mining area boundaries and small-scale features, improving the recognition ability of remote sensing classification for detailed targets. By generating global attention weights, it is possible to capture the overall semantic information of the scene. By weighting the corrected feature, the model can better highlight important features and suppress unimportant features, thus improving the classification accuracy and ensuring the robustness of the model when dealing with complex scenes. The finally obtained fused feature combines information of different modalities, which can effectively improve the performance of scene classification and enhance the robustness and accuracy of the model.

[0103] As Figure 3 shown, the overall process of the mine remote sensing scene classification method according to the embodiment of the present invention includes: The first image data (RGB Image) obtains the first modal feature after being processed by the dual feature enhancement module (DFEM). The first modal feature is processed by two 3×3 convolutional layers to meet the requirements of subsequent convolutional blocks, and then feature extraction is performed through two convolutional blocks (MBConv) and two self-attention mechanism blocks (Self-Attention). Then, global pooling (Global Pool) is used to reduce the spatial dimension of the feature map and extract the global information on each feature channel. Finally, fully connected (FullyConnected, FC) is performed to map the extracted features to the class labels to obtain the first extracted feature. The second image data (SARImage) obtains the second modal feature after being processed by the dual feature enhancement module (DFEM). The second modal feature is processed by two 3×3 convolutional layers to meet the requirements of subsequent convolutional blocks, and then feature extraction is performed through two convolutional blocks (MBConv) and two self-attention mechanism blocks (Self-Attention). Then, global pooling (Global Pool) is used to reduce the spatial dimension of the feature map and extract the global information on each feature channel. Finally, fully connected (Fully Connected, FC) is performed to map the extracted features to the class labels to obtain the second extracted feature. The first extracted feature and the second extracted feature pass through the attention adaptive feature interaction aggregation module (AAFAIM), that is, the feature interaction aggregation module, to obtain the fused feature. The fused feature is subjected to scene classification through a classifier (Classifier).

[0104] As Figure 6 shown, a mine remote sensing scene classification device 600 provided by the embodiment of the present invention includes:

[0105] An acquisition module 610, which is used to acquire the first image data and the second image data corresponding to the target mountain area.

[0106] A feature enhancement module 620, which is used to enhance the boundary features of the first image data through a first feature enhancement network to obtain the first modal feature, and enhance the boundary features of the second image data through a second feature enhancement network to obtain the second modal feature.

[0107] A feature extraction module 630, which is used to extract the first modal feature through a first feature extraction network to obtain the first extracted feature, and extract the second modal feature through a second feature extraction network to obtain the second extracted feature.

[0108] A feature interaction aggregation module 640, which is configured to fuse the first extracted feature and the second extracted feature through a feature interaction aggregation network to obtain a fused feature.

[0109] A classification module 650, which is configured to process the fused feature through a classifier to obtain a scene classification result of the target mountain area.

[0110] As Figure 7 shown, an electronic device 700 provided by an embodiment of the present invention includes a memory 710 and a processor 720; the memory 710 is configured to store a computer program; the processor 720 is configured to implement the above-mentioned mine remote sensing scene classification method when executing the computer program.

[0111] A computer-readable storage medium provided by an embodiment of the present invention has a computer program stored thereon, and when the computer program is executed by a processor, the above-mentioned mine remote sensing scene classification method is implemented.

[0112] Now, an electronic device 700 that can be used as a server or a client of the present invention will be described. It is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device 700 is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device 700 can also represent various forms of mobile devices, such as, personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0113] The electronic device 700 includes a computing unit that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) or a computer program loaded from a storage unit into a random access memory (RAM). In the RAM, various programs and data required for device operation can also be stored. The computing unit, the ROM, and the RAM are connected to each other through a bus. An input / output (I / O) interface is also connected to the bus.

[0114] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc. In this application, the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present invention. In addition, the functional units in various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0115] Although the present invention is disclosed as above, the scope of protection of the present invention is not limited thereto. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention, and these changes and modifications will all fall within the scope of protection of the present invention.

Claims

1. A mine remote sensing scene classification method, characterized in that: include: Acquire first image data and second image data corresponding to the target mountain area; Enhance the boundary features of the first image data by a first feature enhancement network to obtain a first modality feature, and enhance the boundary features of the second image data by a second feature enhancement network to obtain a second modality feature; Extracting the first modal feature through a first feature extraction network to obtain a first extracted feature, and extracting the second modal feature through a second feature extraction network to obtain a second extracted feature; fusing the first extracted features and the second extracted features through a feature interaction aggregation network to obtain a fused feature; The fusion features are processed by a classifier to obtain a scene classification result of the target mountain area.

2. The mine remote sensing scene classification method according to claim 1, characterized in that: The step of enhancing the boundary feature of the first image data by using a first feature enhancement network to obtain a first modality feature, and enhancing the boundary feature of the second image data by using a second feature enhancement network to obtain a second modality feature, includes: Processing the boundary features of the first image data through a first asymmetric convolutional network to obtain a first boundary enhancement feature, and processing the boundary features of the second image data through a second asymmetric convolutional network to obtain a second boundary enhancement feature; The first boundary enhancement feature is processed by a first patch-aware attention mechanism to obtain the first modal feature, and the second boundary enhancement feature is processed by a second patch-aware attention mechanism to obtain the second modal feature, wherein the first feature enhancement network includes the first asymmetric convolutional network and the first patch-aware attention mechanism, and the second feature enhancement network includes the second asymmetric convolutional network and the second patch-aware attention mechanism.

3. The mine remote sensing scene classification method according to claim 2, characterized in that: The first boundary enhancement feature is processed by a first patch-aware attention mechanism to obtain the first modal feature, and the second boundary enhancement feature is processed by a second patch-aware attention mechanism to obtain the second modal feature, including: Extracting the first boundary enhancement feature through a multi-branch feature extraction network to obtain a first multi-branch extraction feature, extracting the second boundary enhancement feature through the multi-branch feature extraction network to obtain a second multi-branch extraction feature, wherein the first patch-aware attention mechanism and the second patch-aware attention mechanism both include the multi-branch feature extraction network and the attention mechanism, the multi-branch feature extraction network processes the first boundary enhancement feature and the second boundary enhancement feature in parallel through local branches, global branches and serial convolution branches, and adjusts the local branches and the global branches through a patch size parameter, and the patch size parameter represents the spatial size of the patch; The first multi-branch extracted features are subjected to feature enhancement processing through the attention mechanism to obtain the first modal features, and the second multi-branch extracted features are subjected to feature enhancement processing through the attention mechanism to obtain the second modal features.

4. The mine remote sensing scene classification method according to claim 1, characterized in that: Before extracting the first modal feature through the first feature extraction network to obtain the first extracted feature, and extracting the second modal feature through the second feature extraction network to obtain the second extracted feature, the method further includes: The convolutional layer and the attention layer are stacked to form a preset stacked network structure, which serve as the first feature extraction network and the second feature extraction network respectively, wherein the structures of the first feature extraction network and the second feature extraction network are the same.

5. The mine remote sensing scene classification method according to claim 4 is characterized in that: The preset stacked network structure includes at least two convolution blocks and at least two self-attention mechanism blocks, wherein the convolution block includes a first convolution layer, a second convolution layer and a third convolution layer, the first convolution layer is used to expand features, the second convolution layer is used to extract features, and the third convolution layer is used to compress features; the self-attention mechanism block is connected to a feedforward neural network, and the feedforward neural network is used to expand and compress the features to be processed.

6. The mine remote sensing scene classification method according to claim 1, characterized in that: The step of fusing the first extracted features and the second extracted features through a feature interaction aggregation network to obtain a fused feature includes: Combining the first extracted feature with the second extracted feature to obtain a combined feature; Performing feature correction on the combined feature by using a correction weight matrix to obtain a correction feature, wherein the correction weight matrix is ​​used to adjust the channel importance weight of the combined feature; The fused feature is obtained through the corrected feature.

7. The mine remote sensing scene classification method according to claim 6, characterized in that: The obtaining of the fusion feature by the correction feature comprises: Splitting the calibration feature into a first calibration feature and a second calibration feature in a channel dimension; Obtaining local attention weights and global attention weights through the correction features; Obtaining a first attention weight by using the local attention weight and the global attention weight; Adding the first attention weight to the first correction feature to obtain a first weighted feature; Adding a second attention weight to the second corrected feature to obtain a second weighted feature; The first weighted feature and the second weighted feature are added to obtain the fused feature, wherein the sum of the first attention weight and the second attention weight is 1.

8. A mine remote sensing scene classification device, characterized in that: include: An acquisition module, the acquisition module is used to acquire first image data and second image data corresponding to the target mountain area; A feature enhancement module, the feature enhancement module is used to enhance the boundary features of the first image data through a first feature enhancement network to obtain a first modal feature, and enhance the boundary features of the second image data through a second feature enhancement network to obtain a second modal feature; A feature extraction module, the feature extraction module is used to extract the first modal feature through a first feature extraction network to obtain a first extracted feature, and extract the second modal feature through a second feature extraction network to obtain a second extracted feature; A feature interaction aggregation module, wherein the feature interaction aggregation module is used to fuse the first extracted feature and the second extracted feature through a feature interaction aggregation network to obtain a fused feature; A classification module is used to process the fusion features through a classifier to obtain a scene classification result of the target mountain area.

9. An electronic device, characterized in that: including memory and processor; The memory is used to store computer programs; The processor is used to implement the mine remote sensing scene classification method as described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the mine remote sensing scene classification method according to any one of claims 1 to 7 is implemented.