Mine land occupation model training method based on feature interaction and scene-semantic collaboration

Through feature interaction and scene-semantic collaboration methods, the performance of the mine land occupancy model in the mine land occupancy recognition task is improved, and the problem of poor performance of the model in pixel-level segmentation tasks is solved, achieving higher learning ability and task performance.

CN119180979BActive Publication Date: 2025-05-13CHINA UNIV OF GEOSCIENCES (WUHAN) +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411001631.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-25
Publication Date
2025-05-13
Estimated Expiration
2044-07-25

AI Technical Summary

Technical Problem

Existing models perform poorly in mine land-based identification classification tasks, especially on pixel-level segmentation tasks, making it difficult to effectively deal with the complex shapes, textures and colors of mines.

Method used

The mining land-occupy model training method based on feature interaction and scene-semantic collaboration is adopted. By obtaining the scene classification data set, the initial training model is used for feature extraction and interaction processing, and task interaction is combined with the self-attention mechanism to generate richer feature representations and more accurate scene classification results.

Benefits of technology

It significantly improves the learning ability and generalization ability of the mine land occupation model, and improves the performance and accuracy of the model in scenario classification and mine land occupation recognition tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119180979B_ABST
    Figure CN119180979B_ABST
Patent Text Reader

Abstract

The present invention provides a method for training a mine land occupation model based on feature interaction and scene-semantic collaboration, and relates to the field of deep learning technology. The method comprises: obtaining a scene classification data set, wherein the scene classification data set comprises a plurality of remote sensing images; inputting the remote sensing image into an initial training model to obtain a prediction result, obtaining a loss value according to the prediction result, and tuning the initial training model by the loss value to obtain a mine land occupation model; the initial training model comprises a feature extraction module and a feature interaction module, and the feature extraction module is used to extract features from the remote sensing image to obtain shallow feature data and deep feature data; the feature interaction module is used to perform task interaction processing on the shallow feature data and the deep feature data through a self-attention mechanism to obtain target feature data, and the target feature data is used to obtain the prediction result. The present invention improves the data feature processing capability of the existing model in the mine land occupation task when facing the differences of different tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning technology, and in particular to a method, device, electronic device and storage medium for training a mine land occupation model based on feature interaction and scene-semantic collaboration. Background Art

[0002] Currently, using large-scale data to train various models has become an effective method, which enables the model to learn more generalized feature representations. This generalized feature is beneficial to downstream tasks and can help the model better understand the common patterns in the input data, thereby improving performance and alleviating overfitting problems. Through the general features learned on large-scale data, the pre-trained model performs better on specific tasks.

[0003] However, the transfer performance of the training model between different tasks varies significantly. Although the use of large-scale data for training can help the model learn generalized feature representations, the differences between different training tasks cause the model to perform poorly on certain specific tasks, limiting the effect of transfer learning and resulting in performance degradation. For example, models pre-trained for scene-level classification usually perform well on similar classification tasks because these models have learned feature representations for different scenes and objects. However, in tasks such as pixel-level segmentation, the model is required not only to understand the different categories in the scene, but also to classify each pixel. Therefore, when faced with the problem of mine land identification and classification, because mines usually have complex shapes, textures, and colors, the model will perform poorly on pixel-level segmentation tasks. Summary of the invention

[0004] The problem solved by the present invention is how to improve the data feature processing capability of the existing model in the mine occupation task when facing the differences of different tasks.

[0005] In order to solve the above problems, the present invention provides a mine occupation model training method, device, electronic device and storage medium based on feature interaction and scene-semantic collaboration.

[0006] In a first aspect, the present invention provides a method for training a mine land occupation model based on feature interaction and scene-semantic collaboration, comprising:

[0007] Acquire a scene classification dataset, wherein the scene classification dataset includes a plurality of remote sensing images;

[0008] Inputting the remote sensing image into an initial training model to obtain a prediction result, obtaining a loss value according to the prediction result, and optimizing the initial training model by using the loss value to obtain a mine land occupation model;

[0009] The initial training model includes a feature extraction module and a feature interaction module.

[0010] Performing feature extraction on the remote sensing image by the feature extraction module to obtain shallow feature data and deep feature data;

[0011] The shallow feature data and the deep feature data are processed through the feature interaction module through a self-attention mechanism to obtain target feature data, and the target feature data is used to obtain the prediction result.

[0012] Optionally, obtaining a scene classification dataset includes:

[0013] Acquire images of the current study area and mine sub-classification annotation data, and obtain a remote sensing multi-modal semantic segmentation dataset based on the images of the current study area and mine sub-classification annotation data;

[0014] The remote sensing multimodal semantic segmentation dataset is expanded and label converted to obtain the scene classification dataset.

[0015] Optionally, the feature extraction module includes a multi-scale attention unit and a pyramid pooling unit, and the feature extraction module is used to extract features from the remote sensing image to obtain shallow feature data and deep feature data, including:

[0016] Performing feature extraction on the remote sensing image to obtain the shallow feature data;

[0017] The shallow feature data is input into the multi-scale attention unit and the pyramid pooling unit respectively to obtain a first deep feature and a second deep feature, wherein the deep feature data includes the first deep feature and the second deep feature.

[0018] Optionally, the feature interaction module includes two task query units and a multi-task feature interaction unit, and the feature interaction module processes the shallow feature data and the deep feature data through a self-attention mechanism to obtain target feature data, and the target feature data is used to obtain the prediction result, including:

[0019] Inputting the first deep feature and the second deep feature into the multi-task feature interaction unit, performing task interaction processing through a self-attention mechanism, and obtaining first deep data and second deep data;

[0020] The shallow feature data are respectively combined with the first deep data and the second deep data, and each is input into the two task query units to perform query processing on tasks of different granularities to obtain a first specified feature and a second specified feature. The target feature data includes the first specified feature and the second specified feature, and the target feature data is used to obtain the prediction result.

[0021] Optionally, the step of inputting the first deep feature and the second deep feature into the multi-task feature interaction unit, performing task interaction processing through a self-attention mechanism, and obtaining the first deep data and the second deep data includes:

[0022] After the first deep features and the second deep features are spliced ​​along the channel, tensor reshaping and layer normalization are performed, and then task interaction integration is performed through multi-head self-attention to obtain the first deep data and the second deep data.

[0023] Optionally, the tasks of different granularities include scene-scale tasks and pixel-scale tasks, and the shallow feature data is respectively combined with the first deep data and the second deep data, and each is input into two task query units to perform query processing on the tasks of different granularities to obtain the first specified feature and the second specified feature, including:

[0024] By means of the scene scale task, the shallow feature data and the first deep data are subjected to scene image classification to obtain the first designated feature;

[0025] Through the pixel scale task, the shallow feature data and the second deep data are semantically segmented to obtain the second specified feature.

[0026] Optionally, obtaining a loss value according to the prediction result, and optimizing the initial training model by using the loss value to obtain a mine land occupation model includes:

[0027] The prediction result and the label data in the remote sensing image are subjected to loss calculation to obtain the loss value, and the initial training model is tuned by the loss value to obtain the mine area model.

[0028] In a second aspect, the present invention provides a mine land occupation model training device based on feature interaction and scene-semantic collaboration, comprising:

[0029] An acquisition unit, configured to acquire a scene classification data set, wherein the scene classification data set includes a plurality of remote sensing images;

[0030] A training unit, used for inputting the remote sensing image into an initial training model to obtain a prediction result, obtaining a loss value according to the prediction result, and optimizing the initial training model by using the loss value to obtain a mine land occupation model;

[0031] The initial training model includes a feature extraction module and a feature interaction module.

[0032] Performing feature extraction on the remote sensing image by the feature extraction module to obtain shallow feature data and deep feature data;

[0033] The shallow feature data and the deep feature data are processed through the feature interaction module through a self-attention mechanism to obtain target feature data, and the target feature data is used to obtain the prediction result.

[0034] In a third aspect, the present invention provides an electronic device, including a memory and a processor;

[0035] The memory is used to store computer programs;

[0036] The processor is used to implement the mine occupation model training method based on feature interaction and scene-semantic collaboration as described in the first aspect when executing the computer program.

[0037] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the mine occupation model training method based on feature interaction and scene-semantic collaboration as described in the first aspect is implemented.

[0038] The beneficial effects of the mine land occupation model training method, device, electronic device and storage medium based on feature interaction and scene-semantic collaboration of the present invention are as follows: the scene classification data set of the present invention includes multiple remote sensing images, and the training model identifies the features in different scenes, providing rich scene information and semantic labels. The remote sensing images are input into the initial training model for training, and the initial feature representation and model parameters can be obtained, laying the foundation for subsequent tuning and optimization. By extracting the shallow features and deep features of the remote sensing image, the model can capture the feature information of the image at different levels and abstract levels; then the shallow and deep features are interactively processed through the self-attention mechanism, which can effectively capture the dependency relationship and semantic information between the features, and improve the model's feature extraction ability for different scene-level classification tasks. The present invention, based on the training method of feature interaction and scene-semantic collaboration, obtains prediction results that will contain richer feature representations and more accurate scene classification results, and the loss function obtained according to the prediction results will better reflect the training effect of the model in the scene classification task, which is helpful to fine-tune the model parameters and optimize the model performance. This solves the problem that when faced with the problem of mine land identification and classification, the model will perform poorly on pixel-level segmentation tasks because mines usually have complex shapes, textures and colors. The training method of the present invention enables the mine land model to better understand image content and semantic relationships, and when faced with the problem of mine land identification and classification, it significantly improves the learning and generalization capabilities of the mine land model, thereby improving the performance and accuracy of the model in scene classification and mine land identification tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 A flowchart of a mine land occupation model training method based on feature interaction and scene-semantic collaboration according to an embodiment of the present invention;

[0040] Figure 2 A schematic diagram of the structure of a mine land occupation model based on feature interaction and scene-semantic collaboration according to an embodiment of the present invention;

[0041] Figure 3 The figure is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0042] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below in conjunction with the accompanying drawings. Although certain embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be interpreted as being limited to the embodiments described herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not intended to limit the scope of protection of the present invention.

[0043] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.

[0044] The term "including" and its variations used in this document are open inclusions, that is, "including but not limited to"; the term "based on" means "based at least in part on"; the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments"; the term "optionally" means "optional embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts of "first", "second", etc. mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0045] It should be noted that the modifications of "one" and "plurality" mentioned in the present invention are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0046] The names of the messages or information exchanged between multiple devices in the embodiments of the present invention are only used for illustrative purposes, and are not used to limit the scope of these messages or information.

[0047] like Figure 1 As shown, a mine land occupation model training method based on feature interaction and scene-semantic collaboration provided by an embodiment of the present invention includes:

[0048] Step S1, obtaining a scene classification dataset, wherein the scene classification dataset includes a plurality of remote sensing images;

[0049] Specifically, this embodiment needs to obtain a scene classification data set including a plurality of remote sensing images. The remote sensing images may be from satellite, airplane or drone photography, and are used to train and verify the mine land occupation model.

[0050] Step S2, inputting the remote sensing image into the initial training model to obtain a prediction result, obtaining a loss value according to the prediction result, and optimizing the initial training model by using the loss value to obtain a mine land occupation model;

[0051] Specifically, this embodiment obtains a loss value from the prediction result, and a cross entropy loss function or other suitable loss function can be used to calculate the loss value. The loss function measures the difference between the output predicted by the model and the actual label. Then, a gradient descent or other optimization algorithm is applied to tune the initial training model to reduce the loss function and improve the accuracy of the model.

[0052] Step S3, wherein the initial training model includes a feature extraction module and a feature interaction module,

[0053] Performing feature extraction on the remote sensing image by the feature extraction module to obtain shallow feature data and deep feature data;

[0054] The shallow feature data and the deep feature data are processed through the feature interaction module through a self-attention mechanism to obtain target feature data, and the target feature data is used to obtain the prediction result.

[0055] Specifically, the feature extraction module of this embodiment can use but is not limited to deep learning models such as convolutional neural networks to extract shallow features and deep features of remote sensing images. These features will be used for subsequent feature interaction processing; the feature interaction module of this embodiment interacts with the deep feature data and the shallow feature data, and realizes the correlation and weight calculation between features through the self-attention mechanism. The self-attention mechanism allows the model to assign different attention weights to different parts of the input when learning feature representation. This embodiment constructs and optimizes the mine land occupation model by interactively processing the multi-layer features of the remote sensing image to achieve accurate classification and identification of the mine land occupation.

[0056] The scene classification data set in this embodiment includes multiple remote sensing images. The training model recognizes the features in different scenes and provides rich scene information and semantic labels. The remote sensing images are input into the initial training model for training to obtain the initial feature representation and model parameters, laying the foundation for subsequent tuning and optimization. By extracting the shallow and deep features of the remote sensing images, the model can capture the feature information of the image at different levels and abstract levels; then the shallow and deep features are interactively processed through the self-attention mechanism, which can effectively capture the dependency and semantic information between the features, and improve the model's feature extraction ability for different scene-level classification tasks. According to the training method of feature interaction and scene-semantic collaboration, the prediction results obtained in this embodiment will contain richer feature representations and more accurate scene classification results, and the loss function obtained according to the prediction results will better reflect the training effect of the model in the scene classification task, which is helpful to fine-tune the model parameters and optimize the model performance. This solves the problem that when faced with the problem of mine land identification and classification, the model will perform poorly on pixel-level segmentation tasks because mines usually have complex shapes, textures and colors. The training method of this embodiment enables the mine land model to better understand image content and semantic relationships. When faced with the problem of mine land identification and classification, it significantly improves the learning and generalization capabilities of the mine land model, thereby improving the performance and accuracy of the model in scene classification and mine land identification tasks.

[0057] Optionally, obtaining a scene classification dataset includes:

[0058] Acquire images of the current study area and mine sub-classification annotation data, and obtain a remote sensing multi-modal semantic segmentation dataset based on the images of the current study area and mine sub-classification annotation data;

[0059] The remote sensing multimodal semantic segmentation dataset is expanded and label converted to obtain the scene classification dataset.

[0060] Specifically, the diffusion model is adopted in this embodiment to expand the remote sensing multimodal semantic segmentation dataset to maintain the diversity of generated data, and the largest number of label categories are calculated through semantic segmentation labels, which are used as scene-scale labels to achieve label conversion processing.

[0061] Optionally, the feature extraction module includes a multi-scale attention unit and a pyramid pooling unit, and the feature extraction module is used to extract features from the remote sensing image to obtain shallow feature data and deep feature data, including:

[0062] Performing feature extraction on the remote sensing image to obtain the shallow feature data;

[0063] The shallow feature data is input into the multi-scale attention unit and the pyramid pooling unit respectively to obtain a first deep feature and a second deep feature, wherein the deep feature data includes the first deep feature and the second deep feature.

[0064] Specifically, shallow feature extraction helps to reduce the dimension and complexity of data. In practical applications, ablation experiments can be added to prove and ensure the effectiveness and relevance of the extracted features to downstream tasks.

[0065] Optionally, the feature interaction module includes two task query units and a multi-task feature interaction unit, and the feature interaction module processes the shallow feature data and the deep feature data through a self-attention mechanism to obtain target feature data, and the target feature data is used to obtain the prediction result, including:

[0066] Inputting the first deep feature and the second deep feature into the multi-task feature interaction unit, performing task interaction processing through a self-attention mechanism, and obtaining first deep data and second deep data;

[0067] The shallow feature data are respectively combined with the first deep data and the second deep data, and each is input into the two task query units to perform query processing on tasks of different granularities to obtain a first specified feature and a second specified feature. The target feature data includes the first specified feature and the second specified feature, and the target feature data is used to obtain the prediction result.

[0068] Specifically, this embodiment uses a multi-task feature interaction unit to perform interactive processing of features between different levels and different tasks. It helps the model learn a richer and more comprehensive feature representation, improves the expressiveness of the features and the generalization ability of the model. The self-attention mechanism allows the model to dynamically assign different attention weights to different parts of the input. This allows the model to assign different weights to different parts of the input when learning feature representations, thereby improving the model's ability to express and understand the input data; this embodiment can realize the query of tasks of different granularities through the processing of the task query unit, further improving the model's ability to learn and express the correlation between different tasks and features.

[0069] Optionally, the step of inputting the first deep feature and the second deep feature into the multi-task feature interaction unit, performing task interaction processing through a self-attention mechanism, and obtaining the first deep data and the second deep data includes:

[0070] After the first deep features and the second deep features are spliced ​​along the channel, tensor reshaping and layer normalization are performed, and then task interaction integration is performed through multi-head self-attention to obtain the first deep data and the second deep data.

[0071] Specifically, this embodiment can merge the features of the two parts through channel splicing, and tensor reshaping can make the model more flexible to adjust the dimension and shape of the data, so that the model can better handle the input multimodal data; layer normalization can help accelerate the convergence speed of the network, improve the training effect, reduce the gradient vanishing or gradient explosion problems that occur during model training, and improve the stability and generalization ability of the model; and then use the multi-head self-attention mechanism, so that the model can integrate the features of different parts under different attention mechanisms, so as to better capture the relationship and mutual influence between the features. This embodiment can better integrate the two deep features through multi-head self-attention, channel splicing, tensor reshaping and layer normalization, improve the expressiveness of the features and the overall learning ability of the model, so that the model can better adapt to complex and changing tasks and improve model performance.

[0072] Optionally, the tasks of different granularities include scene-scale tasks and pixel-scale tasks, and the shallow feature data is respectively combined with the first deep data and the second deep data, and each is input into two task query units to perform query processing on the tasks of different granularities to obtain the first specified feature and the second specified feature, including:

[0073] By means of the scene scale task, the shallow feature data and the first deep data are subjected to scene image classification to obtain the first designated feature;

[0074] Through the pixel scale task, the shallow feature data and the second deep data are semantically segmented to obtain the second specified feature.

[0075] Specifically, this embodiment enables the model to learn feature representations of different tasks at the same time by introducing shallow feature data and deep feature data into scene-scale tasks and pixel-scale tasks respectively, thereby improving the model's adaptability and expressiveness for tasks of different scales. In scene-scale tasks, scene image classification is performed on shallow feature data and deep feature data, which enables the model to learn a more global image feature representation, which helps to classify and understand the overall scene; by inputting shallow feature data and deep feature data into pixel-scale tasks and performing semantic segmentation processing, it is beneficial for the model to understand and extract the semantics of each pixel in the image, thereby obtaining a more detailed and detailed feature representation. The query processing of different scale tasks in this embodiment helps the model learn feature representations of different levels and granularities, improves the model's understanding and expression capabilities of complex scenes, and thus improves the performance and generalization of the model in scene classification tasks.

[0076] Specifically, in this embodiment, scene scale and pixel scale are used to represent tasks of different granularities. Scene classification is a scene scale task, and its purpose is to classify a scene image block; semantic segmentation is a pixel scale task, and its purpose is to classify each pixel in the sample. Among them, the multi-scale attention unit and the pyramid pooling unit are modules that are more effective for scene classification and semantic segmentation, respectively. It can be understood that these two modules are not fixed. In theory, they can be replaced with other modules for scene classification and semantic segmentation. Their function is to extract features from the perspective of different tasks, and to extract effective features for different tasks from the interactive features as queries.

[0077] Optionally, obtaining a loss value according to the prediction result, and optimizing the initial training model by using the loss value to obtain a mine land occupation model includes:

[0078] The prediction result and the label data in the remote sensing image are subjected to loss calculation to obtain the loss value, and the initial training model is optimized by the loss value to obtain the mine area model.

[0079] Specifically, this embodiment compares the features generated by the model with the label data, calculates the loss function, evaluates the difference between the model prediction and the actual label, optimizes the model through the loss function, and performs model tuning based on the loss function to improve the accuracy and generalization of the model. The model will update the parameters according to the feedback of the loss function so that the model better fits the label data, thereby improving the model performance and prediction accuracy. In this embodiment, by calculating the loss of the first specified feature and the second specified feature with the label data, the multimodal data is fully utilized, which helps to better capture the multiple features of the image, improve the understanding and modeling capabilities of complex scenes, and provide an effective training method for the construction of mine land occupation models.

[0080] Combination Figure 2 As shown, in one embodiment, shallow feature extraction is first performed on the remote sensing image to obtain shallow feature data x, and x is respectively input into the edge-enhanced multi-scale attention module (equivalent to the multi-scale attention unit in the embodiment of the present invention) and the pyramid pooling module (equivalent to the pyramid pooling unit in the embodiment of the present invention), and the scene-scale task and the pixel-scale task are respectively performed to obtain the first deep feature x1 and the second deep feature x2, and the first deep feature x1 and the second deep feature x2 are input into the multi-task feature interaction module (equivalent to the multi-task feature interaction unit in the embodiment of the present invention), and the task interaction of each task is captured by the self-attention mechanism, and then the two outputs are respectively combined with the shallow feature data x, and then input into two task query modules (equivalent to the task query unit in the embodiment of the present invention), and the features of tasks of different granularity extracted by different branches are used as queries, and two specified features (head1 and head2) are decomposed from the feature interaction module through self-attention; the decomposed features are input into different task heads to obtain prediction results, and loss calculation is performed to tune the model, so as to obtain the final mine occupation model. Among them, the multi-task feature interaction module captures the task interaction of each task through Concat "concatenation along the channel"; Reshape "reshaping of tensors, such as reshaping a tensor with a shape of (h,w,c) to a shape of (hw,c)"; LayerNorm "layer normalization" and MSA multi-head self-attention. At the same time, Figure 2 The training process shown in the dotted box can be directly used as the backbone network for other similar data after pre-training.

[0081] An embodiment of the present invention provides a mine land occupation model training device based on feature interaction and scene-semantic collaboration, comprising:

[0082] An acquisition unit, configured to acquire a scene classification data set, wherein the scene classification data set includes a plurality of remote sensing images;

[0083] A training unit, used for inputting the remote sensing image into an initial training model to obtain a prediction result, obtaining a loss value according to the prediction result, and optimizing the initial training model by using the loss value to obtain a mine land occupation model;

[0084] The initial training model includes a feature extraction module and a feature interaction module.

[0085] Performing feature extraction on the remote sensing image by the feature extraction module to obtain shallow feature data and deep feature data;

[0086] The shallow feature data and the deep feature data are processed through the feature interaction module through a self-attention mechanism to obtain target feature data, and the target feature data is used to obtain the prediction result.

[0087] The mine occupation model training device based on feature interaction and scene-semantic collaboration of this embodiment is used to implement the mine occupation model training method based on feature interaction and scene-semantic collaboration as described above. Its advantages over the prior art are the same as the advantages of the above-mentioned mine occupation model training method based on feature interaction and scene-semantic collaboration over the prior art, and will not be repeated here.

[0088] like Figure 3 As shown, an electronic device 300 provided in an embodiment of the present invention includes a memory 310 and a processor 320; the memory 310 is used to store a computer program; the processor 320 is used to implement the above-mentioned mine occupation model training method based on feature interaction and scene-semantic collaboration when executing the computer program.

[0089] In other words, an electronic device 300 includes a memory 310 and a processor 320 coupled to the memory 310; the memory 310 is configured to store a computer program; and the processor 320 is configured to perform the following operations when executing the computer program:

[0090] Acquire a scene classification dataset, wherein the scene classification dataset includes a plurality of remote sensing images;

[0091] Inputting the remote sensing image into an initial training model to obtain a prediction result, obtaining a loss value according to the prediction result, and optimizing the initial training model by using the loss value to obtain a mine land occupation model;

[0092] The initial training model includes a feature extraction module and a feature interaction module.

[0093] Performing feature extraction on the remote sensing image by the feature extraction module to obtain shallow feature data and deep feature data;

[0094] The shallow feature data and the deep feature data are processed through the feature interaction module through a self-attention mechanism to obtain target feature data, and the target feature data is used to obtain the prediction result.

[0095] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the mine occupation model training method based on feature interaction and scene-semantic collaboration as described above is implemented.

[0096] In other words, a non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor performs the following operations:

[0097] Acquire a scene classification dataset, wherein the scene classification dataset includes a plurality of remote sensing images;

[0098] Inputting the remote sensing image into an initial training model to obtain a prediction result, obtaining a loss value according to the prediction result, and optimizing the initial training model by using the loss value to obtain a mine land occupation model;

[0099] The initial training model includes a feature extraction module and a feature interaction module.

[0100] Performing feature extraction on the remote sensing image by the feature extraction module to obtain shallow feature data and deep feature data;

[0101] The shallow feature data and the deep feature data are processed through the feature interaction module through a self-attention mechanism to obtain target feature data, and the target feature data is used to obtain the prediction result.

[0102] An electronic device 300 that can be used as a server or client of the present invention will now be described, which is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device 300 is intended to represent various forms of digital electronic computer equipment, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device 300 can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples, and are not intended to limit the implementation of the present invention described and / or required herein.

[0103] The electronic device 300 includes a computing unit, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) or a computer program loaded from a storage unit into a random access memory (RAM). In the RAM, various programs and data required for the operation of the device can also be stored. The computing unit, ROM, and RAM are connected to each other via a bus. An input / output (I / O) interface is also connected to the bus.

[0104] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc. In the present application, the unit described as a separate component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it may be located in one place, or it may be distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment of the present invention. In addition, each functional unit in each embodiment of the present invention can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0105] Although the present invention is disclosed as above, the protection scope of the present invention is not limited thereto. Those skilled in the art may make various changes and modifications without departing from the spirit and scope of the present invention, and these changes and modifications will fall within the protection scope of the present invention.

Claims

1. A mine land occupation model training method based on feature interaction and scene-semantic collaboration, characterized in that: include: Acquire a scene classification dataset, wherein the scene classification dataset includes a plurality of remote sensing images; Inputting the remote sensing image into an initial training model to obtain a prediction result, obtaining a loss value according to the prediction result, and optimizing the initial training model by using the loss value to obtain a mine land occupation model; The initial training model includes a feature extraction module and a feature interaction module. The feature extraction module is used to extract features from the remote sensing image to obtain shallow feature data and deep feature data, wherein the feature extraction module includes a multi-scale attention unit and a pyramid pooling unit; the feature extraction module is used to extract features from the remote sensing image to obtain shallow feature data and deep feature data, specifically including: extracting features from the remote sensing image to obtain the shallow feature data; inputting the shallow feature data into the multi-scale attention unit and the pyramid pooling unit respectively to obtain a first deep feature and a second deep feature, wherein the deep feature data includes the first deep feature and the second deep feature; The shallow feature data and the deep feature data are subjected to task interaction processing through a self-attention mechanism through the feature interaction module to obtain target feature data, and the target feature data is used to obtain the prediction result, wherein the feature interaction module includes two task query units and a multi-task feature interaction unit; the shallow feature data and the deep feature data are subjected to task interaction processing through a self-attention mechanism through the feature interaction module to obtain target feature data, and the target feature data is used to obtain the prediction result, specifically including: inputting the first deep feature and the second deep feature into the multi-task feature interaction unit, performing task interaction processing through a self-attention mechanism, and obtaining first deep data and second deep data; combining the shallow feature data with the first deep data and the second deep data respectively, and inputting each of the two task query units to perform query processing of tasks of different granularities to obtain a first specified feature and a second specified feature, and the target feature data includes the first specified feature and the second specified feature, and the target feature data is used to obtain the prediction result.

2. The method for training a mine land occupation model based on feature interaction and scene-semantic collaboration according to claim 1 is characterized in that: The step of obtaining a scene classification dataset includes: Acquire images of the current study area and mine sub-classification annotation data, and obtain a remote sensing multi-modal semantic segmentation dataset based on the images of the current study area and mine sub-classification annotation data; The remote sensing multimodal semantic segmentation dataset is expanded and label converted to obtain the scene classification dataset.

3. The mine occupation model training method based on feature interaction and scene-semantic collaboration according to claim 1 is characterized in that: The step of inputting the first deep feature and the second deep feature into the multi-task feature interaction unit, performing task interaction processing through a self-attention mechanism, and obtaining first deep data and second deep data includes: After the first deep features and the second deep features are spliced ​​along the channel, tensor reshaping and layer normalization are performed, and then task interaction integration is performed through multi-head self-attention to obtain the first deep data and the second deep data.

4. The method for training a mine land occupation model based on feature interaction and scene-semantic collaboration according to claim 1 is characterized in that: The tasks of different granularities include scene-scale tasks and pixel-scale tasks. The shallow feature data is respectively combined with the first deep data and the second deep data, and each is input into two task query units to perform query processing on the tasks of different granularities to obtain the first specified feature and the second specified feature, including: By means of the scene scale task, the shallow feature data and the first deep data are subjected to scene image classification to obtain the first designated feature; Through the pixel scale task, the shallow feature data and the second deep data are semantically segmented to obtain the second specified feature.

5. The method for training a mine land occupation model based on feature interaction and scene-semantic collaboration according to claim 4 is characterized in that: The step of obtaining a loss value according to the prediction result, and optimizing the initial training model by using the loss value to obtain a mine land occupation model includes: The prediction result and the label data in the remote sensing image are subjected to loss calculation to obtain the loss value, and the initial training model is tuned by the loss value to obtain the mine area model.

6. A mine land occupation model training device based on feature interaction and scene-semantic collaboration, characterized in that: include: An acquisition unit, configured to acquire a scene classification data set, wherein the scene classification data set includes a plurality of remote sensing images; A training unit, used for inputting the remote sensing image into an initial training model to obtain a prediction result, obtaining a loss value according to the prediction result, and optimizing the initial training model by using the loss value to obtain a mine land occupation model; The initial training model includes a feature extraction module and a feature interaction module. The feature extraction module is used to extract features from the remote sensing image to obtain shallow feature data and deep feature data, wherein the feature extraction module includes a multi-scale attention unit and a pyramid pooling unit; the feature extraction module is used to extract features from the remote sensing image to obtain shallow feature data and deep feature data, specifically including: extracting features from the remote sensing image to obtain the shallow feature data; inputting the shallow feature data into the multi-scale attention unit and the pyramid pooling unit respectively to obtain a first deep feature and a second deep feature, wherein the deep feature data includes the first deep feature and the second deep feature; The shallow feature data and the deep feature data are subjected to task interaction processing through a self-attention mechanism through the feature interaction module to obtain target feature data, and the target feature data is used to obtain the prediction result, wherein the feature interaction module includes two task query units and a multi-task feature interaction unit; the shallow feature data and the deep feature data are subjected to task interaction processing through a self-attention mechanism through the feature interaction module to obtain target feature data, and the target feature data is used to obtain the prediction result, specifically including: inputting the first deep feature and the second deep feature into the multi-task feature interaction unit, performing task interaction processing through a self-attention mechanism, and obtaining first deep data and second deep data; combining the shallow feature data with the first deep data and the second deep data respectively, and inputting each of the two task query units to perform query processing of tasks of different granularities to obtain a first specified feature and a second specified feature, and the target feature data includes the first specified feature and the second specified feature, and the target feature data is used to obtain the prediction result.

7. An electronic device, characterized in that: including memory and processor; The memory is used to store computer programs; The processor is used to implement the mine occupation model training method based on feature interaction and scene-semantic collaboration as described in any one of claims 1 to 5 when executing the computer program.

8. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, the method for training a mine occupation model based on feature interaction and scene-semantic collaboration as described in any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Mining area land coverage classification method based on depth feature fusion model

    CN113963262A

  • Target semantic segmentation model training method and generation method, equipment and storage medium

    CN118262104A