Remote sensing image ground feature extraction method and device, equipment and storage medium

By building a lightweight initial geographic feature extraction network, the problem of difficult deployment of visual big models on the drone platform is solved, efficient geographic feature extraction is achieved, and the computing efficiency and accuracy of the drone platform are improved.

CN120544079AActive Publication Date: 2025-08-26XI AN JIAOTONG UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510633183.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-08-26
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

It is difficult for drone platforms to deploy visual models with huge computing resources to extract remote sensing image geographic features, and cannot meet the millisecond response requirements of high-frequency decisions.

Method used

Build an initial geographic feature extraction network, including convolution pre-module, patch embedding module, marking encoding module, encoder module and feature mapping module, and obtain the target geographic feature extraction model through training, reducing the number of model parameters and improving computing efficiency.

Benefits of technology

While improving model accuracy, it reduces computing resource requirements, enhances the applicability of the drone platform, and meets the response needs of high-frequency decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544079A_ABST
    Figure CN120544079A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image ground feature extraction method and device, equipment and a storage medium, and relates to the technical field of image feature extraction. The specific implementation scheme comprises the steps of obtaining a first training set and a second training set; constructing an initial ground feature extraction network, and training the initial ground feature extraction network according to the first training set and the second training set to obtain a target ground feature extraction model; and performing ground feature extraction on an input remote sensing image through the target ground feature extraction model to obtain ground feature features. The problem that a visual large model is difficult to deploy on an unmanned aerial vehicle platform for ground feature extraction in the prior art can be solved, the calculation resource requirement of the model is reduced while the model precision is improved, and the applicability to the unmanned aerial vehicle platform is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image feature extraction, and in particular to a method, device, equipment and storage medium for extracting ground object features from remote sensing images. Background Art

[0002] As an aerial platform, drones can be used in remote sensing, surveying, and mapping industries with high-precision sensors and computing platforms to efficiently collect and process aerial data, offering irreplaceable advantages over traditional manned aircraft. Drone platforms can perform remote sensing image retrieval based on the features of objects captured in aerial remote sensing images, ensuring rapid and accurate positioning over large areas during real-time flight. Furthermore, with the advancement of deep learning, large visual models, such as Transformer, have demonstrated performance advantages in extracting object features.

[0003] However, due to the stringent accuracy and real-time requirements of drones during flight, large, parameter-heavy, and computationally intensive visual models are difficult to deploy on drone platforms for ground feature extraction. For example, while the Transformer's global attention mechanism can effectively capture multi-scale semantic information in remote sensing imagery, drone platforms struggle to meet the massive computational resources required, resulting in an inability to meet the millisecond-level response requirements for high-frequency decision-making.

[0004] Therefore, there is an urgent need for a lightweight remote sensing image feature extraction method that can be applied to UAV platforms. Summary of the Invention

[0005] The embodiments of the present application solve the problem in the prior art that large visual models are difficult to deploy on drone platforms for extracting ground feature extraction by providing a method, device, equipment and storage medium for extracting ground feature from remote sensing images. While improving the model accuracy, it reduces the model's computing resource requirements and improves its applicability to drone platforms.

[0006] In a first aspect, an embodiment of the present application provides a method for extracting ground object features from a remote sensing image, comprising:

[0007] Obtain a first training set and a second training set; wherein the first training set includes multiple sample groups, each sample group includes multiple first remote sensing images and multiple second remote sensing images, the first remote sensing images and the second remote sensing images are cropped based on the same original remote sensing image, and the size of the second remote sensing image is larger than the size of the first remote sensing image; the second training set includes multiple sample pairs, each sample pair includes a third remote sensing image and a fourth remote sensing image, and the fourth remote sensing image is the third remote sensing image after the image of the target object type is masked by a semantic mask; construct an initial object feature extraction network, train the initial object feature extraction network according to the first training set and the second training set, and obtain a target object feature extraction model; perform object extraction on the input remote sensing image through the target object feature extraction model , and obtain the ground feature; wherein, the initial ground feature feature extraction network includes a convolution pre-module, a patch embedding module, a label encoding module, an encoder module and a feature mapping module; the remote sensing image is input into the convolution pre-module for feature extraction and dimensionality reduction processing to reduce the number of model parameters to obtain the target convolution feature; the target convolution feature is input into the patch embedding module for block processing and flattening processing to obtain the first intermediate feature; the first intermediate feature is input into the label encoding module for category labeling and position encoding, and random inactivation processing is performed to obtain the second intermediate feature; the second intermediate feature is input into the encoder module for feature extraction to obtain the third intermediate feature; the third intermediate feature is input into the feature mapping module for layer normalization processing, classification token extraction processing, and feature mapping to obtain the ground feature.

[0008] Furthermore, the encoder module includes a plurality of encoder units connected in sequence; the encoder unit is used to perform layer normalization processing on the features input to the encoder unit and then input the features into the multi-head attention layer to obtain a first feature to be processed, perform random inactivation processing on the first feature to be processed and then add it to the features input to the encoder unit to obtain a second feature to be processed, perform layer normalization processing on the second feature to be processed and then input it into the MLP (Multilayer Perceptron) unit to obtain a third feature to be processed, perform random inactivation processing on the third feature to be processed and then add it to the second feature to be processed to obtain the output feature of the encoder unit and output it.

[0009] Furthermore, the convolution pre-module is used to input the input remote sensing image into the first convolution layer to obtain the first convolution feature, perform batch normalization and activation function on the first convolution feature and then input it into the second convolution layer to obtain the second convolution feature, perform batch normalization and activation function on the second convolution feature and then input it into the third convolution layer to obtain the third convolution feature, and perform batch normalization and activation function on the third convolution feature to obtain the target convolution feature.

[0010] Furthermore, the initial ground feature extraction network is trained according to the first training set and the second training set to obtain a target ground feature extraction model, including:

[0011] The first remote sensing image is input into the preset land feature extraction model and passes through feature centering and activation function to obtain the first feature; the first remote sensing image is input into the initial land feature extraction network and passes through the activation function to obtain the second feature; the second remote sensing image is input into the initial land feature extraction network and passes through the activation function to obtain the third feature; according to the first feature, the second feature and the third feature, based on the first loss function, the parameters of the initial land feature extraction network are updated until the first loss function converges to obtain an intermediate land feature extraction network; the third remote sensing image and the fourth remote sensing image are respectively input into the intermediate land feature extraction network to obtain the fourth feature and the fifth feature respectively; the fourth feature and the fifth feature are respectively input into the fully connected network to obtain the sixth feature and the seventh feature respectively; according to the sixth feature and the seventh feature, based on the second loss function, the parameters of the intermediate land feature extraction network are updated until the second loss function converges to obtain the target land feature extraction model.

[0012] Furthermore, the first loss function is shown in formula (1):

[0013]

[0014] In formula (1), L distill represents the first loss function, H(·) represents the cross entropy loss function, N represents the sum of the number of the first remote sensing images and the second remote sensing images corresponding to the same original remote sensing image, M represents the number of the first remote sensing images corresponding to the same original remote sensing image, and NM represents the number of the second remote sensing images corresponding to the same original remote sensing image. represents the i-th first feature, represents the i-th second feature, represents the jth third feature.

[0015] Furthermore, the second loss function is shown in the following formula (2):

[0016]

[0017] In formula (2), L orth represents the second loss function, z S′ img represents the sixth feature, z S′ mask_img Represents the seventh feature, z S′ img ·z S′ mask_img represents the inner product of the sixth and seventh features, ||z S′img || represents the norm of the sixth feature, ||z S′ mask_img || represents the norm of the seventh feature.

[0018] Furthermore, the fourth remote sensing image is obtained through a masking step, which includes:

[0019] According to the third remote sensing image and the user prompt, based on the SAM (Segment Anything) model, an image of the target object type is obtained; the image of the target object type in the third remote sensing image is blocked to obtain a fourth remote sensing image.

[0020] In a second aspect, an embodiment of the present application provides a device for extracting ground object features from a remote sensing image, comprising:

[0021] An acquisition module is used to acquire a first training set and a second training set; wherein the first training set includes multiple sample groups, each sample group includes multiple first remote sensing images and multiple second remote sensing images, the first remote sensing images and the second remote sensing images are both cropped from the same original remote sensing image, and the size of the second remote sensing image is larger than the size of the first remote sensing image; the second training set includes multiple sample pairs, each sample pair includes a third remote sensing image and a fourth remote sensing image, and the fourth remote sensing image is the third remote sensing image after the image of the target object type is masked by a semantic mask.

[0022] The model construction and training module is used to construct an initial ground feature extraction network, train the initial ground feature extraction network according to the first training set and the second training set, and obtain a target ground feature extraction model.

[0023] The extraction module is used to extract the ground objects from the input remote sensing image through the target ground object feature extraction model to obtain the ground object features.

[0024] Among them, the initial ground feature extraction network includes a convolution pre-module, a patch embedding module, a label encoding module, an encoder module and a feature mapping module; the remote sensing image is input into the convolution pre-module for feature extraction and dimensionality reduction processing to reduce the number of model parameters and obtain the target convolution feature; the target convolution feature is input into the patch embedding module for block processing and flattening processing to obtain the first intermediate feature; the first intermediate feature is input into the label encoding module for category labeling and position encoding, and random inactivation processing is performed to obtain the second intermediate feature; the second intermediate feature is input into the encoder module for feature extraction to obtain the third intermediate feature; the third intermediate feature is input into the feature mapping module for layer normalization processing, classification token extraction processing, and feature mapping to obtain the ground feature feature.

[0025] In a third aspect, an embodiment of the present application provides a device comprising: a processor; a memory for storing processor-executable instructions; and a method for implementing the first aspect or any possible implementation of the first aspect when the processor executes the executable instructions.

[0026] In a fourth aspect, an embodiment of the present application provides a non-volatile computer-readable storage medium, which includes a device for storing a computer program or instruction, and when the computer program or instruction is executed, the method of the first aspect or any possible implementation method of the first aspect is implemented.

[0027] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:

[0028] The embodiment of the present application obtains a first training set and a second training set to construct an initial land feature extraction network including a convolutional pre-module for preliminary feature extraction and dimensionality reduction processing, and an encoder module for better extracting semantic information. The initial land feature extraction network is trained to obtain a target land feature extraction network. The target land feature extraction network is used to extract land features from the input remote sensing image to obtain land feature features. This can improve the model accuracy while reducing the model's computing resource requirements and improve its applicability to UAV platforms. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments of the present application or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0030] Figure 1 A flowchart of a method for extracting ground object features from remote sensing images provided in an embodiment of the present application;

[0031] Figure 2 A schematic diagram of the network structure of the initial ground feature extraction network provided in an embodiment of the present application;

[0032] Figure 3 A schematic diagram of the structure of an encoder unit provided in an embodiment of the present application;

[0033] Figure 4 Another flowchart of the method for extracting ground object features from remote sensing images provided in an embodiment of the present application;

[0034] Figure 5 A schematic diagram of the composition of a remote sensing image feature extraction device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0035] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0036] The following description of some of the technologies involved in the embodiments of this application is provided to facilitate understanding and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for the sake of clarity and conciseness, some descriptions of well-known functions and structures are omitted from the following description.

[0037] As an aerial platform, drones can be used in remote sensing, surveying, and mapping industries with high-precision sensors and computing platforms to efficiently collect and process aerial data, offering irreplaceable advantages over traditional manned aircraft. Drone platforms can perform remote sensing image retrieval based on the features of objects captured in aerial remote sensing images, ensuring rapid and accurate positioning over large areas during real-time flight. Furthermore, with the advancement of deep learning, large visual models, such as Transformer, have demonstrated performance advantages in extracting object features.

[0038] However, due to the stringent accuracy and real-time requirements of drones during flight, large, parameter-heavy, and computationally intensive visual models are difficult to deploy on drone platforms for ground feature extraction. For example, while the Transformer's global attention mechanism can effectively capture multi-scale semantic information in remote sensing imagery, drone platforms struggle to meet the massive computational resources required, resulting in an inability to meet the millisecond-level response requirements for high-frequency decision-making.

[0039] Therefore, there is an urgent need for a lightweight remote sensing image feature extraction method that can be applied to UAV platforms.

[0040] Against this background technology, the present disclosure provides a method for extracting ground object features from remote sensing images, which can solve the problem in the existing technology that large visual models are difficult to deploy on drone platforms for ground object feature extraction. While improving the model accuracy, it reduces the model's computing resource requirements and improves its applicability to drone platforms.

[0041] The execution subject of the remote sensing image feature extraction method provided by the embodiment of the present disclosure can be a computer or a server, or can also be other electronic devices with data processing capabilities; alternatively, the execution subject of the method can also be a processor (such as a central processing unit (CPU)) in the above electronic device; alternatively, the execution subject of the method can also be an application (application, APP) installed in the above electronic device that can implement the functions of the method; alternatively, the execution subject of the method can also be a functional module or unit in the above electronic device that has the functions of the method. There is no limitation on the execution subject of the method herein.

[0042] The following is an exemplary description of the method for extracting ground object features from remote sensing images with reference to the accompanying drawings.

[0043] Figure 1 : is a flow chart of the method for extracting ground feature from remote sensing images provided in the embodiment of the present application. Figure 1 This is only an execution order shown in the embodiment of the present application, and does not represent the only execution order of the remote sensing image feature extraction method. Figure 1 The steps shown can be performed in parallel or in reverse. Figure 1 As shown, the method may include S101 to S103.

[0044] S101: Obtain a first training set and a second training set.

[0045] Among them, the first training set includes multiple sample groups, each sample group includes multiple first remote sensing images and multiple second remote sensing images, the first remote sensing images and the second remote sensing images are both cropped based on the same original remote sensing image, and the size of the second remote sensing image is larger than that of the first remote sensing image; the second training set includes multiple sample pairs, each sample pair includes a third remote sensing image and a fourth remote sensing image, and the fourth remote sensing image is the third remote sensing image after the image of the target object type is masked by a semantic mask.

[0046] Exemplarily, the first remote sensing image may be a global view image including most areas in the original remote sensing image, and the second remote sensing image may be a local view image including a small area in the original remote sensing image.

[0047] For example, taking the original remote sensing images including image A, image B, and image C as an example, image A, image B, and image C are processed in sequence. By cropping image A separately, we can obtain a global view image A1 including most areas of image A, a global view image A2 including most areas of image A, a local view image a1 including a small area of ​​image A, a local view image a2 including a small area of ​​image A, and a local view image a3 including a small area of ​​image A; by cropping image B separately, we can obtain a global view image B1 including most areas of image B, a global view image B2 including most areas of image B, a local view image b1 including a small area of ​​image B, a local view image b2 including a small area of ​​image B, and a local view image b3 including a small area of ​​image B; by cropping image C separately, we can obtain a global view image C1 including most areas of image C, a global view image C2 including most areas of image C, a local view image c1 including a small area of ​​image C, a local view image c2 including a small area of ​​image C, and a local view image c3 including a small area of ​​image C.

[0048] Furthermore, images A1, A2, B1, B2, C1, and C2 can be determined as first remote sensing images, and images a1, a2, a3, b1, b2, b3, c1, c2, and c3 can be determined as second remote sensing images; images A1, A2, a1, a2, and a3 can be taken as a sample group, images B1, B2, b1, b2, and b3 can be taken as a sample group, and images C1, C2, c1, c2, and c3 can be taken as a sample group.

[0049] Exemplarily, the target object types may include roads, buildings, etc., without limitation; the fourth remote sensing image may be obtained by manually masking the image of the target object type in the third remote sensing image through a semantic mask, and different target object types correspond to different semantic masks.

[0050] Illustratively, the target ground object types corresponding to the fourth remote sensing images in the multiple sample pairs may be the same or different.

[0051] For example, the third remote sensing image includes images D1, D2, and D3, and the corresponding fourth remote sensing images are images d1, d2, and d3, respectively. The target object type corresponding to image d1 can be a road, the target object type corresponding to image d2 can be a building, and the target object type corresponding to image d3 can be vegetation.

[0052] S102: Construct an initial ground feature extraction network, and train the initial ground feature extraction network according to the first training set and the second training set to obtain a target ground feature extraction model.

[0053] Among them, the initial ground feature extraction network includes a convolution pre-module, a patch embedding module, a label encoding module, an encoder module and a feature mapping module; the remote sensing image is input into the convolution pre-module for feature extraction and dimensionality reduction processing to reduce the number of model parameters and obtain the target convolution feature; the target convolution feature is input into the patch embedding module for block processing and flattening processing to obtain the first intermediate feature; the first intermediate feature is input into the label encoding module for category labeling and position encoding, and random inactivation processing is performed to obtain the second intermediate feature; the second intermediate feature is input into the encoder module for feature extraction to obtain the third intermediate feature; the third intermediate feature is input into the feature mapping module for layer normalization processing, classification token extraction processing, and feature mapping to obtain the ground feature feature.

[0054] Specifically, the encoder module includes a plurality of encoder units connected in sequence; the encoder unit is used to perform layer normalization processing on the features input to the encoder unit and then input the features into the multi-head attention layer to obtain a first feature to be processed, perform random inactivation processing on the first feature to be processed and then add it to the features input to the encoder unit to obtain a second feature to be processed, perform layer normalization processing on the second feature to be processed and then input it into the MLP unit to obtain a third feature to be processed, perform random inactivation processing on the third feature to be processed and then add it to the second feature to be processed to obtain the output feature of the encoder unit and output it.

[0055] Specifically, the convolution pre-module is used to input the input remote sensing image into the first convolution layer to obtain the first convolution feature, perform batch normalization and activation function on the first convolution feature and then input it into the second convolution layer to obtain the second convolution feature, perform batch normalization and activation function on the second convolution feature and then input it into the third convolution layer to obtain the third convolution feature, and perform batch normalization and activation function on the third convolution feature to obtain the target convolution feature.

[0056] For example, the data of the first training set and the second training set can be directly used to directly train the initial ground feature extraction network through the gradient descent method to obtain the target ground feature extraction model.

[0057] Exemplarily, when performing network training on the initial feature extraction network, a cross entropy loss function and / or an orthogonal loss function may be used as the loss function, without any restriction on the loss function.

[0058] The following uses a specific embodiment to further illustrate the network structure of the initial ground feature extraction network.

[0059] Figure 2 This is a schematic diagram of the network structure of the initial ground feature extraction network provided in the embodiment of the present application. Figure 2 , the convolutional pre-module in the initial ground feature extraction network (i.e. Figure 2The Conv2d in Figure 1 contains three convolutional layers, each followed by a batch normalization layer and a ReLU activation function (the specific structure of the convolutional pre-module is not shown). The first convolutional layer in the convolutional pre-module has a 7×7 kernel with a stride of 2 and a padding of 3; the second convolutional layer has a 3×3 kernel with a stride of 2 and a padding of 1; and the third convolutional layer has a 3×3 kernel with a stride of 4 and a padding of 3. For a remote sensing image of size 224×224×3 input to the convolutional pre-module, a target convolutional feature of size 14×14×384 is obtained.

[0060] Compared with directly using the original image as the input of the entire Transformer network, the number of visual tokens sent to the subsequent Transformer network is significantly reduced, effectively reducing the complexity of subsequent self-attention calculations; at the same time, the local receptive field characteristics of convolution help to extract low-level features such as texture and edges in remote sensing images, providing more semantically informative input for subsequent Transformers; compared with the standard ViT (Vision Transformer) network, by setting up a convolutional front network, the total number of parameters can be reduced by about 30%, while improving feature extraction performance.

[0061] The target convolutional features of size 14×14×384 are input into the patch embedding module in the initial feature extraction network (i.e. Figure 2 After block processing and flattening, the first intermediate feature with a size of 196×384 can be obtained.

[0062] The patch embedding module can convert the target convolutional features into serialized patch tokens and flatten the spatial dimensions through the Flatten operation.

[0063] The first intermediate feature of size 196×384 is input to the label encoding module (i.e. Figure 2 The tag encoding module in the ) performs category tagging and position encoding, and performs random inactivation processing. The tag encoding module first concatenates the first intermediate feature with the classification token (Class token) of size 1×384 (i.e. Figure 2 Concat in ), get the feature of size 197×384, and then combine the feature of size 197×384 with the position code of size 197×384 (i.e. Figure 2 The PositionEmbedding in the , is added to obtain a feature of size 197 × 384, and then the feature of size 197 × 384 is randomly deactivated (i.e. Figure 2 After Dropout in , the second intermediate feature with a size of 197×384 is obtained.

[0064] The second intermediate feature with a size of 197×384 is input into the encoder module (i.e. Figure 2 The TransformerEncoder in

[15] is used for feature extraction to obtain the third intermediate feature with a size of 197×384.

[0065] The encoder module includes 12 encoder units connected in sequence (i.e. Figure 2 Encoder Block in . Figure 3 This is a schematic diagram of the structure of the encoder unit provided in the embodiment of the present application. Figure 3 As shown, each encoder unit performs layer normalization on the features of the input encoder unit (i.e. Figure 3 After Layer Norm in , input the multi-head attention layer with the number of heads number_heads of 6 (i.e. Figure 2 Multi-Head Attention in ), obtain the first feature to be processed; perform random inactivation on the first feature to be processed (i.e. Figure 3 Dropout in ) and then add the feature of the input encoder unit to obtain the second feature to be processed; the second feature to be processed is layer normalized (i.e. Figure 3 Layer Norm in) and then input into the MLP unit (i.e. Figure 3 The third feature to be processed is obtained by performing random inactivation on the third feature to be processed (i.e. Figure 3 After the Dropout in , it is added to the second feature to be processed to obtain the output feature of the encoder unit and output it.

[0066] For example, an MLP unit includes two linear layers. The input features of the MLP unit are activated by an activation function and randomly dropped out after passing through the first linear layer. Then, they are randomly dropped out after passing through the second linear layer to obtain the output features of the MLP unit. The activation function in the MLP unit can be a ReLU activation function or a GELU activation function.

[0067] Through the self-attention mechanism, the model can capture the global relationship between different features, extract semantic information in complex scenes, directly model the long-distance dependency between different objects, and capture the complex spatial semantic structure in remote sensing images, such as road networks, building layouts and other macroscopic surface features.

[0068] The third intermediate feature with a size of 197×384 is input into the feature mapping module (i.e. Figure 2The MLP Head in the layer normalization process, classification token extraction process, and feature mapping are performed. The feature mapping module first performs layer normalization on the third intermediate feature with a size of 197×384 (i.e. Figure 2 The Layer Norm in the , gets the feature of size 197 × 384, and then extracts the classification tokens of the feature of size 197 × 384 (i.e. Figure 2 Extract Class token in ), get the feature of size 1×384, and then extract the feature representation corresponding to the classification token (i.e. Figure 2 Pre-Logits in , and then through the linear layer (i.e. Figure 2 The Linear in is used to map the features to the required output dimension, and finally obtain the ground feature with a size of 1×384.

[0069] S103 , extracting ground objects from the input remote sensing image using a target ground object feature extraction model to obtain ground object features.

[0070] For example, after obtaining the target object feature extraction model, the remote sensing image that needs to be subjected to object feature extraction can be input into the target object feature extraction model. The target object feature extraction model identifies and extracts the target object features in the remote sensing image to obtain the object features corresponding to the remote sensing image.

[0071] The embodiment of the present application obtains a first training set and a second training set to construct an initial land feature extraction network including a convolutional pre-module for preliminary feature extraction and dimensionality reduction processing, and an encoder module for better extracting semantic information. The initial land feature extraction network is trained to obtain a target land feature extraction network. The target land feature extraction network is used to extract land features from the input remote sensing image to obtain land feature features. This can improve the model accuracy while reducing the model's computing resource requirements and improve its applicability to UAV platforms.

[0072] Figure 4 Another flow chart of the method for extracting ground feature from remote sensing images provided in the embodiment of the present application. In some possible implementations, such as Figure 4 As shown, the initial ground object feature extraction network is trained according to the first training set and the second training set to obtain the target ground object feature extraction model, including S401 to S405.

[0073] S401. Input the first remote sensing image into a preset land feature extraction model and pass it through feature centering and activation function to obtain a first feature; input the first remote sensing image into an initial land feature extraction network and pass it through activation function to obtain a second feature; input the second remote sensing image into the initial land feature extraction network and pass it through activation function to obtain a third feature.

[0074] Exemplarily, the preset terrain feature extraction model may be a pre-trained large ViT model (eg, a ViT-Large model) having terrain feature extraction functionality, and the specific structure of the preset terrain feature extraction model is not limited.

[0075] Exemplarily, the activation functions in S401 may all be softmax functions.

[0076] S402. According to the first feature, the second feature, and the third feature, based on the first loss function, update the parameters of the initial ground feature extraction network until the first loss function converges to obtain an intermediate ground feature extraction network.

[0077] For example, during training, the model parameters of the preset ground feature extraction model may be frozen, and only the parameters of the initial ground feature extraction network may be updated.

[0078] Specifically, the first loss function is shown in formula (1):

[0079]

[0080] In formula (1), L distill represents the first loss function, H(·) represents the cross entropy loss function, N represents the sum of the number of the first remote sensing images and the second remote sensing images corresponding to the same original remote sensing image, M represents the number of the first remote sensing images corresponding to the same original remote sensing image, and NM represents the number of the second remote sensing images corresponding to the same original remote sensing image. represents the i-th first feature, represents the i-th second feature, represents the jth third feature.

[0081] It can be understood that by taking the feature distribution output by the preset feature extraction model as the learning target of the initial feature extraction model, and using the cross-entropy loss of the first and second features (both corresponding to the global view), the initial feature extraction network can be forced to approximate the feature distribution of the preset feature extraction model at the overall semantic level, ensuring high-level semantic consistency; using the cross-entropy loss of the first and third features (corresponding to the global view and the local view), the sensitivity of the initial feature extraction network to local texture features can be improved through fine-grained feature matching. This hierarchical loss combination allows the resulting intermediate feature extraction network to inherit the strong representation prior of the preset feature extraction model while also having lightweight architecture characteristics, ultimately achieving a balance between feature extraction accuracy and computational efficiency.

[0082] S403: Input the third remote sensing image and the fourth remote sensing image into the intermediate ground feature extraction network respectively to obtain the fourth feature and the fifth feature respectively.

[0083] S404: Input the fourth feature and the fifth feature into a fully connected network to obtain the sixth feature and the seventh feature respectively.

[0084] It can be understood that the fully connected network is used to further project the features obtained by the intermediate feature extraction network.

[0085] S405 . According to the sixth feature and the seventh feature, based on the second loss function, update the parameters of the intermediate ground feature extraction network until the second loss function converges to obtain the target ground feature extraction model.

[0086] Specifically, the second loss function is shown in the following formula (2):

[0087]

[0088] In formula (2), L orth represents the second loss function, z S′ img represents the sixth feature, z S′ mask_img Represents the seventh feature, z S′ img ·z S′ mask_img represents the inner product of the sixth and seventh features, ||z S′ img || represents the norm of the sixth feature, ||z S′ mask_img || represents the norm of the seventh feature.

[0089] It can be understood that by setting the second loss function, the intermediate feature extraction network can pay more attention to learning the information of the target feature type of the occluded part, that is, the features of roads, buildings and other types of features, and encourage the intermediate feature extraction network to generate feature space distributions as orthogonal as possible for different versions of the same image, thereby improving the ability to discriminate important semantic features of the occluded part.

[0090] This embodiment trains the initial land feature feature extraction network based on the first loss function according to the first feature, the second feature, and the third feature, and obtains an intermediate land feature feature extraction network that inherits the strong representation prior of the preset land feature feature extraction model and has lightweight architecture characteristics. According to the sixth feature and the seventh feature, the intermediate land feature feature extraction network is trained based on the second loss function, and a target land feature feature extraction model that better extracts the semantic information of the land feature of the target land feature type and has higher land feature extraction accuracy can be obtained.

[0091] In some possible implementations, the fourth remote sensing image is obtained through a masking step, and the masking step includes:

[0092] According to the third remote sensing image and the user prompt, an image of the target object type is obtained based on the SAM model; the image of the target object type in the third remote sensing image is blocked to obtain a fourth remote sensing image.

[0093] It can be understood that the SAM model is a pre-trained image segmentation model released by Meta in 2023, which can achieve accurate segmentation with a small amount of prompts (i.e. user prompts).

[0094] For example, the user prompt can be a point prompt or a box prompt. An interactive segmentation tool based on the SAM model can be used to extract an image of the target feature type by inputting the third remote sensing image and the user prompt. A high-precision semantic mask can be generated for the image of the target feature type, ensuring the accuracy of the mask boundary and semantic consistency. The third remote sensing image containing the semantic mask is then determined as the fourth remote sensing image. This allows for the rapid and accurate acquisition of the fourth remote sensing image.

[0095] In some possible implementations, after obtaining the training set, the remote sensing images in the training set can be normalized, and at least one of the hue, saturation, and brightness of the random remote sensing images can be adjusted. Furthermore, the remote sensing images can be translated, scaled, rotated, aspect-ratio-converted, horizontally flipped, and vertically flipped. This can accelerate model training, enhance model robustness, and prevent model overfitting.

[0096] Verification experiment

[0097] Comparative experiment 1: Use different neural network models trained based on the same training set to extract ground feature, and perform remote sensing image retrieval based on the extracted ground feature to compare the accuracy of different neural network models.

[0098] Based on publicly available raw remote sensing images, we constructed a first training set of 18,000 remote sensing images and a second training set of 8,000 remote sensing images. We used different neural network models trained on these first and second training sets (including the Dino-v2(s) model, the Dino-v2(b) model, the ResNet-50 model, the ResNet-101 model, the Ours model (the target feature extraction model of this application), and the Ours' model (a model with additional parameters based on the target feature extraction model of this application)) to extract feature information. Remote sensing image retrieval was performed based on the extracted feature information. The accuracy comparison results are shown in Table 1.

[0099] Table 1 Accuracy comparison results of different neural network models

[0100]

[0101] Among them, SR@k represents the retrieval success rate, which indicates the proportion of query images that contain at least one correct match in the Top-k retrieval results; mAP represents the mean average precision, which is used to measure model performance.

[0102] Referring to the accuracy comparison results in Table 1, it can be seen that the target ground object feature extraction model of the present application outperforms existing mainstream methods in all evaluation indicators. On the one hand, compared with the Dino-v2(s) model and Dino-v2(b) model based on the Vision Transformer model, the target ground object feature extraction model of the present application has greater stability in remote sensing image retrieval tasks while significantly reducing the model size, indicating that the target ground object feature extraction model of the present application has better ground object feature extraction capabilities, especially further expanding its leading advantage in the Top-10 task. On the other hand, compared with the ResNet-50 model and ResNet-101 model, the target ground object feature extraction model of the present application has significantly improved feature expression ability and retrieval performance. On the other hand, generally speaking, the larger the number of model parameters, the better its ability to learn feature details, thus having a greater performance advantage. However, compared with the Ours' model with an increased number of parameters, the target ground object feature extraction model of the present application has better model performance despite having fewer parameters. This shows that the target ground object feature extraction model of the present application has improved performance while having fewer parameters and is more suitable for UAV platforms.

[0103] Comparative experiment 2: The target ground feature extraction model of this application obtained by training based on different training sets is used to extract ground feature, and remote sensing image retrieval is performed based on the extracted ground feature, and the corresponding accuracies of different training sets are compared.

[0104] Based on the first training set and the second training set used in comparative experiment 1, training set A, training set B, training set C, training set D, training set E, and training set F were constructed.

[0105] The difference between training set A, training set B, and training set C is that the target object type corresponding to the fourth remote sensing image in training set A only includes roads (i.e., the occluded area only includes roads). The target object type corresponding to the fourth remote sensing image in training set B only includes buildings (i.e., the occluded area only includes buildings). The target object type corresponding to the fourth remote sensing image in training set C includes both roads and buildings (i.e., the occluded area includes both roads and buildings).

[0106] The difference between training set D and training set A is that in training set D, the fourth remote sensing image of training set A no longer uses semantic mask to block the road area, but only randomly adjusts the brightness and contrast corresponding to the road area to weaken the road area.

[0107] The difference between training set E and training set B is that in training set E, the fourth remote sensing image of training set B no longer uses semantic mask to block the building area, but only randomly adjusts the brightness and contrast corresponding to the building area to weaken the building area.

[0108] The difference between training set F and training set C is that in training set F, the fourth remote sensing image of training set C no longer uses semantic masks to block the road area and building area. Instead, the brightness and contrast corresponding to the road area and building area are randomly adjusted to weaken the road area and building area.

[0109] The brightness can be reduced by 30% and the contrast can be reduced by 20% based on the original value.

[0110] The accuracy comparison results corresponding to different training sets are shown in Table 2.

[0111] Table 2 Comparison of accuracy results corresponding to different training sets

[0112]

[0113] Among them, SR@k represents the retrieval success rate, which indicates the proportion of query images that contain at least one correct match in the Top-k retrieval results; mAP represents the mean average precision, which is used to measure model performance.

[0114] Referring to the accuracy comparison results in Table 2, we can see that training set C performs best across all metrics. Compared to the other training sets, occluding multi-category semantic regions (roads + buildings) significantly improves retrieval success rate and ranking quality, demonstrating that this global semantic occlusion strategy significantly enhances model robustness. Occluding a single category (such as "road" or "building"), namely training sets A and B, performs second best, with the accuracy corresponding to training set A being superior to that of training set B, indicating that building semantic regions contribute more to accuracy.

[0115] Training set F performs worse than training set C, but better than training sets D and E. This shows that the strategy of weakening multi-category semantic regions (roads + buildings) can also effectively improve retrieval accuracy, but the overall effect is not as good as the occlusion strategy. Among the single-category weakening strategies, training sets D and E perform similarly, with the SR@k corresponding to training set E slightly higher than the SR@k corresponding to training set D. However, training set E is slightly lower than training set D in terms of mAP.

[0116] Overall, the strategy of blocking semantic regions outperforms the strategy of weakening them. This suggests that directly blocking semantic regions can more effectively force the model to focus on other key information, thereby improving retrieval performance. For a single semantic category, blocking buildings is more effective than weakening them, while blocking roads is slightly less effective than weakening them. This suggests that different landmark categories contribute differently to retrieval performance, and that architectural semantic regions may be more salient.

[0117] Table 3 shows the parameter count and single-image running time comparison of the target object feature extraction model of the present application and the mainstream feature extraction models (i.e., the various models in comparative experiment 1). Referring to Table 3, it can be seen that the target object feature extraction model of the present application (i.e., Ours model) significantly reduces the parameter count while ensuring feature extraction capabilities, reflecting good lightweight characteristics. Compared with the Dino-v2(s) model and Dino-v2(b) model based on Vision Transformer, the target object feature extraction model of the present application has a smaller feature dimension, reduces storage overhead and computational complexity, and has a faster inference speed, making it suitable for resource-constrained application scenarios. Compared with the ResNet-50 model and ResNet-101 model, the target object feature extraction model of the present application has an advantage in running time, further verifying its effectiveness in improving computational efficiency. Overall, the target object feature extraction model of the present application achieves a balance between computational cost, storage requirements, and inference efficiency in the task of object feature extraction, providing an optimal solution for efficient remote sensing image analysis.

[0118] Table 3 Retrieval time of different feature extraction networks

[0119] Method Parameter quantity Feature Dimension Single running time (seconds) Dino-v2(s) 21M 384 0.19 Dino-v2(b) 86M 768 0.4 ResNet-50 25.6M 2048 0.2 ResNet-101 44.5M 2048 0.29 Ours 18.6M 128 0.15 Ours' 39.8M 256 0.27

[0120] Although the present application provides method operation steps such as embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative work. The order of steps listed in this embodiment is only one way of executing the order of many steps and does not represent the only execution order. When an actual device or client product is executed, it can be executed in the order of the method shown in this embodiment or the accompanying drawings or in parallel (for example, in a parallel processor or multi-threaded processing environment).

[0121] like Figure 5 As shown, the embodiment of the present application also provides a device for extracting ground feature from remote sensing images. The device includes:

[0122] The acquisition module 501 is used to obtain a first training set and a second training set; wherein the first training set includes multiple sample groups, each sample group includes multiple first remote sensing images and multiple second remote sensing images, the first remote sensing images and the second remote sensing images are both cropped from the same original remote sensing image, and the size of the second remote sensing image is larger than the size of the first remote sensing image; the second training set includes multiple sample pairs, each sample pair includes a third remote sensing image and a fourth remote sensing image, and the fourth remote sensing image is the third remote sensing image after the image of the target object type is masked by a semantic mask.

[0123] The model construction and training module 502 is used to construct an initial land feature extraction network, train the initial land feature extraction network according to the first training set and the second training set, and obtain a target land feature extraction model; wherein, the initial land feature extraction network includes a convolution pre-module, a patch embedding module, a label encoding module, an encoder module and a feature mapping module; the remote sensing image is input into the convolution pre-module for feature extraction and dimensionality reduction processing to reduce the number of model parameters to obtain the target convolution feature; the target convolution feature is input into the patch embedding module for block processing and flattening processing to obtain the first intermediate feature; the first intermediate feature is input into the label encoding module for category labeling and position encoding, and random inactivation processing is performed to obtain the second intermediate feature; the second intermediate feature is input into the encoder module for feature extraction to obtain the third intermediate feature; the third intermediate feature is input into the feature mapping module for layer normalization processing, classification token extraction processing, and feature mapping to obtain the land feature feature.

[0124] The extraction module 503 is used to extract the ground objects from the input remote sensing image by using the target ground object feature extraction model to obtain ground object features.

[0125] The beneficial effects and specific implementation methods of the present device embodiment can be referred to the aforementioned method embodiment, and will not be described in detail here.

[0126] Some modules in the apparatus described herein may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, classes, etc. that perform specific tasks or implement specific abstract data types. The present application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0127] The devices or modules described in the above application embodiments can be implemented by computer chips or physical devices, or by products with certain functions. For ease of description, the above devices are described separately by function in various modules. When implementing the embodiments of this application, the functions of each module can be implemented in the same or multiple software and / or hardware. Of course, a module that implements a certain function can also be implemented by combining multiple sub-modules or sub-units.

[0128] The methods, devices, or modules described in this application can be implemented in the form of computer-readable program code. The controller can be implemented in any appropriate manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (such as software or firmware) that can be executed by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to implement the same function of the controller in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered as structures within the hardware component. Or even, the means for realizing various functions may be considered to be both a software module for realizing the method and a structure within a hardware component.

[0129] An embodiment of the present application further provides a device comprising: a processor; a memory for storing processor-executable instructions; and when the processor executes the executable instructions, the method described in the embodiment of the present application is implemented.

[0130] The embodiments of the present application also provide a non-volatile computer-readable storage medium having a computer program or instruction stored thereon. When the computer program or instruction is executed, the method described in the embodiments of the present application is implemented.

[0131] In addition, each functional module in each embodiment of the present invention may be integrated into one processing module, or each module may exist independently, or two or more modules may be integrated into one module.

[0132] The above-mentioned storage medium includes, but is not limited to, random access memory (RAM), read-only memory (ROM), cache, hard disk drive (HDD), or memory card. The memory can be used to store computer program instructions.

[0133] It can be seen from the description of the above implementation methods that those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary hardware. Based on this understanding, the technical solution of the present application can essentially or the part that contributes to the prior art can be embodied in the form of a software product, or it can be embodied through the implementation process of data migration. The computer software product can be stored in a storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a mobile terminal, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present application or certain parts of the embodiments.

[0134] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referenced to each other. Each embodiment focuses on the differences from other embodiments. All or part of this application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, mobile communication terminals, multi-processor systems, microprocessor-based systems, programmable electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc.

[0135] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit the present application. Although the present application has been described in detail with reference to the aforementioned embodiments, a person of ordinary skill in the art should understand that the technical solutions described in the aforementioned embodiments can still be modified, or some or all of the technical features therein can be replaced by equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the present application.

Claims

1. A method for extracting ground feature from remote sensing images, characterized in that: include: Obtain the first training set and the second training set; The first training set includes multiple sample groups, each of which includes multiple first remote sensing images and multiple second remote sensing images. The first remote sensing images and the second remote sensing images are cropped from the same original remote sensing image, and the size of the second remote sensing images is larger than that of the first remote sensing images. The second training set includes multiple sample pairs, each of which includes a third remote sensing image and a fourth remote sensing image. The fourth remote sensing image is the third remote sensing image after the image of the target object type is masked by a semantic mask. Constructing an initial ground feature extraction network, and training the initial ground feature extraction network according to the first training set and the second training set to obtain a target ground feature extraction model; Extracting objects from the input remote sensing image using the target object feature extraction model to obtain object features; The initial ground feature extraction network includes a convolutional pre-module, a patch embedding module, a label encoding module, an encoder module and a feature mapping module; The remote sensing image is input into the convolution pre-module for feature extraction and dimensionality reduction processing to reduce the number of model parameters and obtain the target convolution feature; the target convolution feature is input into the patch embedding module for block processing and flattening processing to obtain the first intermediate feature; the first intermediate feature is input into the label encoding module for category labeling and position encoding, and random inactivation processing is performed to obtain the second intermediate feature; the second intermediate feature is input into the encoder module for feature extraction to obtain the third intermediate feature; the third intermediate feature is input into the feature mapping module for layer normalization processing, classification token extraction processing, and feature mapping to obtain the ground feature.

2. The method according to claim 1, characterized in that The encoder module includes a plurality of encoder units connected in sequence; The encoder unit is used to perform layer normalization processing on the features input to the encoder unit and then input the features into the multi-head attention layer to obtain a first feature to be processed, perform random inactivation processing on the first feature to be processed and add it to the features input to the encoder unit to obtain a second feature to be processed, perform layer normalization processing on the second feature to be processed and then input it into the MLP unit to obtain a third feature to be processed, perform random inactivation processing on the third feature to be processed and add it to the second feature to be processed to obtain the output feature of the encoder unit and output it.

3. The method according to claim 1, characterized in that The convolution pre-module is used to input the input remote sensing image into the first convolution layer to obtain the first convolution feature, perform batch normalization and activation function on the first convolution feature and then input it into the second convolution layer to obtain the second convolution feature, perform batch normalization and activation function on the second convolution feature and then input it into the third convolution layer to obtain the third convolution feature, and perform batch normalization and activation function on the third convolution feature to obtain the target convolution feature.

4. The method according to claim 1, wherein The initial ground feature extraction network is trained according to the first training set and the second training set to obtain a target ground feature extraction model, including: Inputting the first remote sensing image into a preset ground feature extraction model and subjecting it to feature centering and activation function to obtain a first feature; inputting the first remote sensing image into the initial ground feature extraction network and subjecting it to activation function to obtain a second feature; inputting the second remote sensing image into the initial ground feature extraction network and subjecting it to activation function to obtain a third feature; According to the first feature, the second feature, and the third feature, based on a first loss function, updating parameters of the initial ground feature extraction network until the first loss function converges, thereby obtaining an intermediate ground feature extraction network; Inputting the third remote sensing image and the fourth remote sensing image into the intermediate feature extraction network respectively to obtain a fourth feature and a fifth feature respectively; Inputting the fourth feature and the fifth feature into a fully connected network respectively to obtain a sixth feature and a seventh feature respectively; According to the sixth feature and the seventh feature, based on the second loss function, the parameters of the intermediate ground feature extraction network are updated until the second loss function converges, thereby obtaining a target ground feature extraction model.

5. The method according to claim 4, characterized in that The first loss function is shown in formula (1): In formula (1), L distill represents the first loss function, H(·) represents the cross entropy loss function, N represents the sum of the number of first remote sensing images and second remote sensing images corresponding to the same original remote sensing image, M represents the number of first remote sensing images corresponding to the same original remote sensing image, and NM represents the number of second remote sensing images corresponding to the same original remote sensing image. represents the i-th first feature, represents the i-th second feature, represents the jth third feature.

6. The method according to claim 4, characterized in that The second loss function is shown in the following formula (2): In formula (2), L orth represents the second loss function, z S′ img represents the sixth feature, z S′ mask_img Represents the seventh feature, z S′ img ·z S′ mask_img represents the inner product of the sixth and seventh features, ||z S′ img ‖ represents the norm of the sixth feature, ‖z S′ mask_img || represents the norm of the seventh feature.

7. The method according to claim 1, characterized in that The fourth remote sensing image is obtained through a shielding step, and the shielding step includes: Obtaining an image of the target object type based on the SAM model according to the third remote sensing image and the user prompt; The image of the target object type in the third remote sensing image is blocked to obtain the fourth remote sensing image.

8. A device for extracting ground feature from remote sensing images, characterized in that: include: An acquisition module is configured to acquire a first training set and a second training set; wherein the first training set includes a plurality of sample groups, each of which includes a plurality of first remote sensing images and a plurality of second remote sensing images, wherein the first remote sensing images and the second remote sensing images are cropped from the same original remote sensing image, and the size of the second remote sensing images is larger than the size of the first remote sensing images; and the second training set includes a plurality of sample pairs, each of which includes a third remote sensing image and a fourth remote sensing image, wherein the fourth remote sensing image is the third remote sensing image after the image of the target object type is masked by a semantic mask. A model construction and training module is used to construct an initial ground feature extraction network, train the initial ground feature extraction network according to the first training set and the second training set, and obtain a target ground feature extraction model; An extraction module is used to extract ground objects from the input remote sensing image using the target ground object feature extraction model to obtain ground object features; Among them, the initial ground feature extraction network includes a convolution pre-module, a patch embedding module, a label encoding module, an encoder module and a feature mapping module; the remote sensing image is input into the convolution pre-module for feature extraction and dimensionality reduction processing to reduce the amount of model parameters and obtain the target convolution feature; the target convolution feature is input into the patch embedding module for block processing and flattening processing to obtain the first intermediate feature; the first intermediate feature is input into the label encoding module for category labeling and position encoding, and random inactivation processing is performed to obtain the second intermediate feature; the second intermediate feature is input into the encoder module for feature extraction to obtain the third intermediate feature; the third intermediate feature is input into the feature mapping module for layer normalization processing, classification token extraction processing, and feature mapping to obtain the ground feature feature.

9. A device for performing a method for extracting features of ground objects from remote sensing images, characterized in that: include: processor; a memory for storing processor-executable instructions; When the processor executes the executable instructions, the method according to any one of claims 1 to 7 is implemented.

10. A non-volatile computer-readable storage medium, characterized in that: The device comprises a computer program or an instruction for storing the computer program or the instruction, which, when executed, enables the method according to any one of claims 1 to 7 to be implemented.

Citation Information

Patent Citations

  • High-resolution remote sensing image classification method based on feature pooling and divisive normalization expression

    CN106845417A

  • Method, system, device and medium for processing surface feature image

    CN117473027A

  • Enhanced semantic-position feature fusion network method and device based on different pre-training feature extraction backbone and application

    CN118196628A

  • High-resolution remote sensing image road extraction method, device and equipment and storage medium

    CN119399637A

  • Remote sensing image ground object segmentation method based on hierarchical feature extraction and decoding

    CN119672555A