Ground feature element extraction method and device based on prior embedding, equipment and medium
By introducing attention-based feature extraction and convolutional feature extraction networks with regional prior information, and combining them with feature fusion networks, the problem of spatiotemporal heterogeneity of ground features in optical satellite imagery is solved, and high-precision ground feature extraction is achieved.
Patent Information
- Application Number
- CN202511656313.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-12
- Publication Date
- 2026-03-20
AI Technical Summary
Existing methods for extracting ground features cannot accurately extract features from optical satellite imagery due to the spatiotemporal heterogeneity of these features, resulting in low extraction accuracy and failing to meet the requirements of high-precision applications.
A prior embedding-based method for extracting ground features is adopted. By combining an attention feature extraction network and a convolutional feature extraction network, regional prior information is introduced to generate global image features. The extracted ground features are then generated through a feature fusion network and trained using a preset loss function.
It improves the model's generalization ability and robustness, enhances the accuracy of ground feature extraction, and meets the needs of high-precision applications.
Smart Images

Figure CN121708490A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of remote sensing image processing technology, and in particular to a method, apparatus, device and medium for extracting ground features based on prior embedding. Background Technology
[0002] As the basic units of natural and artificial cover on the Earth's surface, the spatial distribution and dynamic changes of ground features can directly reflect the resource allocation, environmental evolution and intensity of human activities in a specific area. Accurately extracting ground features from optical satellite images is not only the core task of remote sensing image information interpretation, but also an important foundation for supporting planning decisions. It has important practical significance in scenarios such as land spatial planning, ecological protection area delineation, and urban construction.
[0003] Existing methods for extracting ground features can be broadly categorized into two types: traditional machine learning methods and deep learning methods. Traditional machine learning methods rely on expert knowledge to design identifying features for ground features, such as spectral features, texture features, and shape features. The machine learning model is then trained and used to perform inference based on these features. Deep learning methods, on the other hand, automatically extract features from optical satellite imagery using end-to-end deep learning models. This eliminates the need for manually designed identifying features, and its core advantage lies in its ability to adaptively extract features and model complex scenes.
[0004] However, existing algorithms assume that similar land cover features have stable characteristics in optical satellite imagery, focusing on mining features from massive amounts of data and extracting land cover features based on the image features of optical satellite imagery. These algorithms are insufficiently adapted to the spatiotemporal heterogeneity of the land surface and cannot directly address land cover features with strong spatiotemporal heterogeneity. For example, factors such as illumination, phenology, and regional culture can lead to inconsistent or even conflicting identification features for the same type of land cover feature at different times or in different regions. This can cause existing models that rely solely on image feature recognition to make misjudgments, making it difficult to accurately extract land cover features from optical satellite imagery.
[0005] In summary, existing methods for extracting ground features suffer from low accuracy, making it difficult to meet the application requirements for high-precision ground feature extraction. Summary of the Invention
[0006] This invention provides a method, apparatus, device, and medium for extracting ground features based on prior embedding, which addresses the shortcomings of existing ground feature extraction methods, such as low accuracy and difficulty in meeting the application requirements of high-precision ground feature extraction.
[0007] This invention provides a method for extracting ground features based on prior embedding, comprising: acquiring optical satellite imagery and regional prior information of the optical satellite imagery; the optical satellite imagery is obtained by taking pictures of the target area, and the regional prior information is information characterizing the regional features of the target area; inputting the optical satellite imagery and the regional prior information into a ground feature extraction model to obtain the ground feature extraction results of the optical satellite imagery output by the ground feature extraction model; wherein, the ground feature extraction model is trained based on sample optical satellite imagery, sample regional prior information corresponding to the sample optical satellite imagery, and sample ground feature extraction results corresponding to the sample optical satellite imagery.
[0008] According to the present invention, a method for extracting ground features based on prior embedding is provided. The ground feature extraction model includes an attention feature extraction network, a convolutional feature extraction network, a feature pyramid network, and a feature fusion network. The attention feature extraction network, the feature pyramid network, and the feature fusion network are connected sequentially, and the convolutional feature extraction network and the feature fusion network are connected. Specifically, the convolutional feature extraction network is used to extract spatial features and local semantic features from optical satellite imagery to generate local image features; the attention feature extraction network and the feature pyramid network are used to embed prior information and extract global semantic features from optical satellite imagery and regional prior information to generate global image features; and the feature fusion network is used to generate ground feature extraction results based on local image features and global image features.
[0009] According to the prior embedding-based method for extracting ground features provided by the present invention, the attention feature extraction network includes a Stem module, a first stage module, a second stage module, a third stage module, and a fourth stage module connected in sequence, and the feature pyramid network includes a first network layer, a second network layer, a third network layer, and a fourth network layer; wherein, the first stage module, the second stage module, the third stage module, and the fourth stage module all include a PEBlock module, and the PEBlock module includes a prior embedding window attention module; the first stage module is connected to the first network layer, the second stage module is connected to the second network layer, the third stage module is connected to the third network layer, and the fourth stage module is connected to the fourth network layer.
[0010] According to the present invention, a method for extracting ground features based on prior embedding is provided. The global image features include a first global feature map, a second global feature map, a third global feature map, and a fourth global feature map. Specifically, a Stem module is used to extract features from optical satellite imagery to generate an initial feature map; a first Stage module is used to embed prior information and extract global semantic features from the initial feature map and regional prior information to generate the first feature map; a first network layer is used to perform convolution and upsampling processing on the first feature map to generate the first global feature map; and a second Stage module is used to embed prior information and extract global semantic features from the first feature map and regional prior information. The system performs the following steps: Global semantic feature extraction to generate a second feature map; a second network layer performs convolution and upsampling on the second feature map to generate a second global feature map; a third stage module embeds prior information from the second feature map and region prior information and extracts global semantic features to generate a third feature map; a third network layer performs convolution and upsampling on the third feature map to generate a third global feature map; a fourth stage module embeds prior information from the third feature map and region prior information and extracts global semantic features to generate a fourth feature map; a fourth network layer performs convolution on the fourth feature map to generate a fourth global feature map.
[0011] According to the prior embedding-based method for extracting ground features provided by the present invention, the feature fusion network includes a fifth network layer and a sixth network layer; wherein, the fifth network layer is used to perform feature fusion on the first global feature map, the second global feature map, the third global feature map, the fourth global feature map and local image features to generate a fused feature map; the sixth network layer is used to perform upsampling processing on the fused feature map to generate the ground feature extraction result.
[0012] According to the prior embedding-based method for extracting ground features provided by the present invention, the ground feature extraction model is trained based on a preset loss function; wherein, the preset loss function is determined based on geometric constraint loss and edge loss, and the edge loss is determined based on cross-entropy dice joint loss, binary cross-entropy loss and focal loss.
[0013] According to the prior embedding-based method for extracting ground features provided by the present invention, the ground feature extraction model is trained based on the following steps: acquiring sample optical satellite images, prior information of sample areas corresponding to the sample optical satellite images, and sample ground feature extraction results corresponding to the sample optical satellite images; training an initial PEDNet model based on the sample optical satellite images, prior information of sample areas, and sample ground feature extraction results to obtain the ground feature extraction model.
[0014] This invention also provides a ground feature extraction device based on prior embedding, comprising: an acquisition module for acquiring optical satellite imagery and regional prior information of the optical satellite imagery; the optical satellite imagery is obtained by photographing the target area, and the regional prior information is information characterizing the regional features of the target area; and a ground feature extraction module for inputting the optical satellite imagery and the regional prior information into a ground feature extraction model to obtain the ground feature extraction results of the optical satellite imagery output by the ground feature extraction model; wherein the ground feature extraction model is trained based on sample optical satellite imagery, sample regional prior information corresponding to the sample optical satellite imagery, and sample ground feature extraction results corresponding to the sample optical satellite imagery.
[0015] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement any of the above-described methods for extracting ground features based on prior embedding.
[0016] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described methods for extracting ground features based on prior embedding.
[0017] The method, apparatus, equipment, and medium for extracting ground features based on prior embedding provided by this invention fully consider the potential conflict of image features that may exist for the same type of ground features in different regions or under different imaging conditions. Regional prior information is introduced during the ground feature extraction process. This regional prior information is information that is distinct from image features but can characterize the regional features of the target area. After inputting the optical satellite image of the target area and the regional prior information into the ground feature extraction model, the model can simultaneously identify and extract ground features in the target area based on the image features of the optical satellite image and the regional features in the regional prior information. This avoids the model's misjudgment problem that may result from relying solely on image features, improves the model's generalization ability and robustness, and thus improves the accuracy of ground feature extraction, thereby meeting the application requirements for high-precision ground feature extraction. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating the method for extracting geographic features based on prior embedding provided by the present invention.
[0020] Figure 2 This is a schematic diagram of the structure of the feature extraction model provided by the present invention.
[0021] Figure 3 This is a structural schematic diagram of the PEWA module provided by the present invention.
[0022] Figure 4 This is a schematic diagram of a typical area image slice and mask provided by the present invention.
[0023] Figure 5 This is one of the schematic diagrams showing the image comparison of ablation experimental test results provided by the present invention.
[0024] Figure 6 This is the second schematic diagram comparing the image results of the ablation experiment provided by the present invention.
[0025] Figure 7 This is a schematic diagram of the structure of the ground feature extraction device based on prior embedding provided by the present invention.
[0026] Figure 8 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0028] Please see Figures 1 to 6 , Figure 1 This is a flowchart illustrating the ground feature extraction method based on prior embedding provided by the present invention. Figure 2 This is a schematic diagram of the structure of the feature extraction model provided by the present invention. Figure 3 This is a schematic diagram of the PEWA module provided by the present invention. Figure 4 This is a schematic diagram of a typical region image slice and mask provided by the present invention. Figure 5 This is one of the schematic diagrams comparing the image results of the ablation experiment provided by the present invention. Figure 6 This is the second schematic diagram comparing the image results of the ablation experiment provided by the present invention.
[0029] like Figure 1 As shown, in this embodiment, the method for extracting ground features based on prior embedding includes steps S110 to S120, each of which is detailed below: S110: Acquire optical satellite imagery and regional prior information of optical satellite imagery.
[0030] Optical satellite imagery is obtained by taking pictures of a target area, and the prior information of the area is information that characterizes the features of the ground features in the target area.
[0031] S120: Input optical satellite imagery and regional prior information into the ground feature extraction model to obtain the ground feature extraction results of the optical satellite imagery output by the ground feature extraction model.
[0032] The feature extraction model is trained based on sample optical satellite imagery, prior information of the sample area corresponding to the sample optical satellite imagery, and the feature extraction results of the sample optical satellite imagery.
[0033] Specifically, this embodiment addresses the problem of insufficient spatiotemporal heterogeneity adaptation in existing mainstream deep learning models. Based on the existing mainstream deep learning model framework, it proposes a deep network model (PEDNet) with dual-path collaboration of convolutional network branches and attention mechanism network branches based on prior knowledge embedding.
[0034] During the model training phase, sample optical satellite images, prior information of the sample area corresponding to the sample optical satellite images, and sample ground feature extraction results corresponding to the sample optical satellite images can be obtained first.
[0035] Furthermore, the initial PEDNet model is trained using sample optical satellite imagery, prior information of the sample area corresponding to the sample optical satellite imagery, and the sample feature extraction results corresponding to the sample optical satellite imagery to obtain the feature extraction model.
[0036] Furthermore, in the model inference (i.e. model application) stage, the optical satellite imagery of the target area and the regional prior information used to characterize the regional features of the target area are input into the trained feature extraction model to obtain the feature extraction results of the optical satellite imagery.
[0037] Specifically, such as Figure 2As shown, the encoder part of the ground feature extraction model using the PEDNet model framework contains two network branches: one is a feature extraction path based on an attention mechanism, namely the attention feature extraction network, which focuses on realizing global semantic feature extraction based on prior knowledge embedding. The attention feature extraction network can capture regional correlation features and global image features through the PEWA (PriorEmbedded Window Attention) module, which can effectively model long-distance feature correlations across regions and scenes and adapt to feature changes of ground features in different spatiotemporal scenarios; the other is a feature extraction path based on a convolutional neural network, namely the convolutional feature extraction network, which relies on the local receptive field advantage of the convolutional neural network to focus on extracting fine-grained spatial details and local semantic information in optical satellite imagery. The two network branches work together to represent the features of ground features from three levels: regional prior, global image features, and local image features. Finally, the decoder part fuses and infers the features of different network branches to generate the ground feature extraction results.
[0038] In particular, to address the potential conflict of image features in optical satellite imagery across different scenarios, the feature extraction model can vector-encode the prior information (i.e., prior knowledge) used to characterize the features of the land cover region and embed it into the attention feature extraction network to construct a collaborative representation of the features of the land cover image and the features of the land cover region. This strengthens the model's encoding of differentiated regional features of the same or similar land cover in optical satellite imagery and improves the model's feature discrimination ability and generalization in complex scenarios.
[0039] Optionally, the optical satellite imagery used to input the ground feature extraction model can adopt a uniform image size as needed, such as 512×512 pixel optical satellite imagery.
[0040] Optionally, regional prior information can be manually encoded into easily readable string data with actual physical meaning, so that the feature extraction model can directly read it and perform vector encoding through the PEWA module.
[0041] Optionally, the vectorized embedding of region prior information can be independent of the attention feature extraction network, forming a separate network branch, which is then fused with the convolutional feature extraction network and the attention feature extraction network for decoding and inference.
[0042] Optionally, the prior information for the region includes, but is not limited to, imaging condition information of optical satellite imagery, geospatial information of the target region (such as longitude, latitude, topography, etc.), seasonal information, climate information, and regional cultural information of the target region.
[0043] Optionally, regional prior information can be any information that can be accurately and efficiently expressed through expert experience, is related to the characteristics of the target region, and is largely independent of image features.
[0044] The prior embedding-based ground feature extraction method provided in this embodiment fully considers the potential conflict of image features for the same type of ground features in different regions or under different imaging conditions. It introduces regional prior information into the ground feature extraction process. This regional prior information is information that distinguishes itself from image features but can characterize the regional features of the target area. After inputting the optical satellite image of the target area and the regional prior information into the ground feature extraction model, the model can simultaneously identify and extract ground features in the target area based on both the image features of the optical satellite image and the regional features in the regional prior information. This avoids the model's misjudgment problem that may result from relying solely on image features, improves the model's generalization ability and robustness, and thus increases the accuracy of ground feature extraction, thereby meeting the application requirements for high-precision ground feature extraction.
[0045] In some embodiments, the ground feature extraction model includes an attention feature extraction network, a convolutional feature extraction network, a feature pyramid network, and a feature fusion network. The attention feature extraction network, the feature pyramid network, and the feature fusion network are connected sequentially, and the convolutional feature extraction network and the feature fusion network are connected. The convolutional feature extraction network is used to extract spatial features and local semantic features from optical satellite imagery to generate local image features. The attention feature extraction network and the feature pyramid network are used to embed prior information and extract global semantic features from optical satellite imagery and regional prior information to generate global image features. The feature fusion network is used to generate ground feature extraction results based on local and global image features.
[0046] Please continue reading. Figure 2 In this embodiment, the ground feature extraction model using the PEDNet model framework includes an attention feature extraction network, a convolutional feature extraction network, a feature pyramid network (FPN), and a feature fusion network. The attention feature extraction network, the feature pyramid network, and the feature fusion network are connected in sequence, and the convolutional feature extraction network and the feature fusion network are connected.
[0047] Specifically, after inputting the optical satellite imagery of the target area and the regional prior information used to characterize the regional features of the target area into the feature extraction model, the convolutional feature extraction network can extract spatial features and local semantic features from the optical satellite imagery, generate local image features, and input the local image features into the feature fusion network.
[0048] Meanwhile, the attention feature extraction network and the feature pyramid network can embed prior information and extract global semantic features from optical satellite imagery and regional prior information to generate global image features, and input the global image features into the feature fusion network.
[0049] Furthermore, the feature fusion network can perform feature fusion and upsampling on local and global image features to generate the final ground feature extraction results.
[0050] In some embodiments, the attention feature extraction network includes a Stem module, a first stage module, a second stage module, a third stage module, and a fourth stage module connected in sequence, and the feature pyramid network includes a first network layer, a second network layer, a third network layer, and a fourth network layer; wherein, the first stage module, the second stage module, the third stage module, and the fourth stage module all include a PEBlock module, and the PEBlock module includes a prior embedding window attention module; the first stage module is connected to the first network layer, the second stage module is connected to the second network layer, the third stage module is connected to the third network layer, and the fourth stage module is connected to the fourth network layer.
[0051] like Figure 2 As shown, the attention feature extraction network includes a Stem module, a first stage module (denoted as Stage1), a second stage module (denoted as Stage2), a third stage module (denoted as Stage3), and a fourth stage module (denoted as Stage4) connected in sequence. The feature pyramid network includes a first network layer, a second network layer, a third network layer, and a fourth network layer. The first stage module is connected to the first network layer, the second stage module is connected to the second network layer, the third stage module is connected to the third network layer, and the fourth stage module is connected to the fourth network layer.
[0052] The first stage module, the second stage module, the third stage module, and the fourth stage module all include the PEBlock module, which includes the Prior Embedded Window Attention (PEWA) module.
[0053] The Stem module can perform preliminary feature extraction on the input optical satellite imagery through two convolution operations to generate an initial feature map.
[0054] The first, second, third, and fourth stage modules can gradually refine features based on the initial feature map through multiple layers of PEBlock modules, generating more detailed feature maps.
[0055] Optionally, the second, third, and fourth stage modules can also perform dimensionality reduction and dimensionality increase of the feature maps through the PatchMerging operation to meet the needs of subsequent feature extraction and feature fusion.
[0056] In some embodiments, the global image features include a first global feature map, a second global feature map, a third global feature map, and a fourth global feature map; wherein, the Stem module is used to extract features from the optical satellite image to generate an initial feature map; the first Stage module is used to embed prior information and extract global semantic features from the initial feature map and regional prior information to generate a first feature map; a first network layer is used to perform convolution and upsampling processing on the first feature map to generate a first global feature map; and the second Stage module is used to embed prior information and extract global semantic features from the first feature map and regional prior information. The process involves generating a second feature map; a second network layer performing convolution and upsampling on the second feature map to generate a second global feature map; a third stage module embedding prior information and extracting global semantic features from the second feature map and region prior information to generate a third feature map; a third network layer performing convolution and upsampling on the third feature map to generate a third global feature map; a fourth stage module embedding prior information and extracting global semantic features from the third feature map and region prior information to generate a fourth feature map; and a fourth network layer performing convolution on the fourth feature map to generate a fourth global feature map.
[0057] like Figure 2 and Figure 3 As shown, the core unit of the PEBlock module is the PEWA module. The design of the PEWA module closely matches the core needs of extracting ground features from cross-scene optical satellite imagery, and its advantages are particularly prominent in complex scenarios. Traditional attention mechanism network branches often suffer from feature confusion when processing optical satellite imagery because they ignore the regional characteristics of the target area (such as geographical differences, different imaging conditions, etc.). However, the PEWA module constructs a dual learnable feature system that combines "regional prior features" and "image features" by embedding bias encoding with prior knowledge, enabling the model to specifically enhance the discriminative power of ground features.
[0058] It should be noted that, from the perspective of the implementation mechanism of the PEWA module, the core process of the PEWA module mainly includes feature map generation, prior bias fusion, similarity calculation and attention output.
[0059] First, regarding the feature map of the input PEWA module... ( Represents the matrix dimension. For the number of channels, For feature map height, For feature map The PEWA module can generate a query matrix Q, a key matrix K, and a value matrix V through a 1×1 convolutional network, which serve as the basis for subsequent attention calculations. The calculation formulas for the query matrix Q, key matrix K, and value matrix V are as follows: ; ; ; ; in, Represents a convolutional network; matrix dimensions and feature maps of the query matrix Q, key matrix K, and value matrix V. Maintaining consistency ensures complete preservation of the input feature map. Spatial structure and channel information.
[0060] Secondly, in order to integrate regional prior information into global image features, the PEWA module can generate prior codes through the PEBlock module. Here, the regional prior information can be split into at least one prior code by the PEBlock module according to actual needs. For example, geographic location information and imaging time information can be encoded independently, and discrete regions can be identified based on the prior codes. The process described above by the PEBlock module, which projects the embedding layer and a 3×3 convolutional network onto the dimension of the matching feature map, can be expressed by the following formula: .
[0061] Furthermore, the PEWA module can use a 16×16 window to divide the feature map, balancing the scale coverage of ground features and background interference in the feature map.
[0062] Optionally, to address the issues in traditional self-attention mechanisms ( The computational bottleneck (the number of pixels in the window) can be approximated by the softmax function using L2 normalization and Taylor expansion, linearizing the weight calculation. This allows the calculation of the first (number of) pixels within the window to be optimized. query vectors With the Key vectors Similarity between It can be expressed by the following formula: ; in, Represent the L2 norm to ensure Taylor approximation The effectiveness.
[0063] The attention output within the window is a weighted sum of the value features, calculated as follows: ; Finally, the PEWA module can perform attention modulation and output spatiotemporal attention. The calculation formula for spatiotemporal attention is as follows: ; The above formula, through the continuous effect of prior coding, supports the model's ability to accurately extract ground features from cross-space and cross-temporal optical satellite images within a large window range, and enhances the model's ability to learn the features of ground features in complex spatiotemporal scenarios.
[0064] Specifically, from the overall process of the attention feature extraction network and the feature pyramid network, the Stem module can perform preliminary feature extraction on the input optical satellite imagery and regional prior information through two convolution operations to generate an initial feature map.
[0065] For the initial feature map, the first stage module can perform prior information embedding and global semantic feature extraction on the initial feature map and region prior information. The PEWA module in the first stage module can sequentially perform the above-mentioned feature mapping generation, prior bias fusion, similarity calculation and attention output on the initial feature map and region prior information, and finally generate the first feature map.
[0066] The output of the first Stage module can be used as the input of the first network layer in the feature pyramid network. The first network layer can perform convolution and upsampling on the first feature map to generate the first global feature map.
[0067] Furthermore, the output of the first stage module can also be used as the input of the second stage module. The second stage module can perform prior information embedding and global semantic feature extraction on the first feature map and region prior information. The PEWA module in the second stage module can sequentially perform the above-mentioned feature mapping generation, prior bias fusion, similarity calculation and attention output on the first feature map and region prior information, and finally generate the second feature map.
[0068] The output of the second Stage module can be used as the input to the second network layer in the feature pyramid network. The second network layer can perform convolution and upsampling on the second feature map to generate the second global feature map.
[0069] Similarly, the output of the second stage module can also be used as the input of the third stage module. The third stage module can perform prior information embedding and global semantic feature extraction on the second feature map and region prior information. The PEWA module in the third stage module can perform the above-mentioned feature mapping generation, prior bias fusion, similarity calculation and attention output on the second feature map and region prior information in sequence, and finally generate the third feature map.
[0070] The output of the third Stage module can be used as the input to the third network layer in the feature pyramid network. The third network layer can perform convolution and upsampling on the third feature map to generate the third global feature map.
[0071] Similarly, the output of the third stage module can also be used as the input of the fourth stage module. The fourth stage module can perform prior information embedding and global semantic feature extraction on the third feature map and region prior information. The PEWA module in the fourth stage module can perform the above-mentioned feature mapping generation, prior bias fusion, similarity calculation and attention output on the third feature map and region prior information in sequence, and finally generate the fourth feature map.
[0072] The output of the fourth Stage module can be used as the input to the fourth network layer in the feature pyramid network. The fourth network layer can perform convolution processing on the fourth feature map to generate the fourth global feature map.
[0073] Among them, the first global feature map, the second global feature map, the third global feature map, and the fourth global feature map are global image features at different scales.
[0074] In some embodiments, the feature fusion network includes a fifth network layer and a sixth network layer; wherein, the fifth network layer is used to perform feature fusion on the first global feature map, the second global feature map, the third global feature map, the fourth global feature map and local image features to generate a fused feature map; the sixth network layer is used to perform upsampling processing on the fused feature map to generate ground feature extraction results.
[0075] like Figure 2 As shown, the feature fusion network includes a fifth network layer and a sixth network layer. The fifth network layer can perform feature fusion on the first global feature map, the second global feature map, the third global feature map, and the fourth global feature map output by the attention feature extraction network and the feature pyramid network, as well as the local image features output by the convolutional feature extraction network, to generate a fused feature map.
[0076] Furthermore, the sixth network layer can upsample the fused feature map to generate the final ground feature extraction result.
[0077] In some embodiments, the feature extraction model is trained based on a preset loss function; wherein the preset loss function is determined based on geometric constraint loss and edge loss, and the edge loss is determined based on cross-entropy dice joint loss, binary cross-entropy loss and focal loss.
[0078] Understandably, while optimizing the model framework of the ground feature extraction model, in order to address the challenges in the task of extracting ground features from optical satellite imagery, such as the easy loss of edge details, the difficulty in maintaining geometric structure, and the uneven distribution of different ground feature types, this embodiment designs a loss function that is more suitable for extracting ground features from optical satellite imagery.
[0079] The preset loss functions include Geometry Constraint Loss and Edge Loss. Edge Loss can be further decomposed into Cross Entropy Dice Joint Loss, Binary Cross Entropy Loss, and Focal Loss. By fusing multiple losses and adding them together with certain weights, the advantages of different loss functions can be comprehensively utilized to guide the model to focus more on a few categories of ground features.
[0080] Among them, the preset loss function The expression is as follows: ; in, For geometric constraint loss, This is the edge loss.
[0081] The expression for edge loss is as follows: ; in, The cross-entropy dice joint loss, For binary cross-entropy loss, The loss is the focus.
[0082] In some embodiments, the feature extraction model is trained based on the following steps: acquiring sample optical satellite images, prior information of the sample area corresponding to the sample optical satellite images, and the feature extraction results of the sample optical satellite images; and training the initial PEDNet model based on the sample optical satellite images, prior information of the sample area, and the feature extraction results of the sample optical satellite images to obtain the feature extraction model.
[0083] This embodiment presents a ground feature extraction method based on prior embedding. Building upon the mainstream CNN and Transformer dual-path deep model framework, it proposes a PEDNet model that utilizes a dual-path collaboration between a convolutional network branch and an attention mechanism network branch based on prior knowledge embedding. Specifically, addressing the potential conflict of ground feature features in optical satellite imagery across different scenarios, this embodiment vectorizes the regional prior information used to characterize the ground feature area of the target region and embeds it into the attention mechanism network branch. This constructs a collaborative representation of ground feature image features and regional features, strengthening the model's differential feature encoding for ground features with identical or similar image features, and improving the model's feature discrimination ability and generalization in complex scenarios.
[0084] Compared with existing technologies, the prior embedding-based ground feature extraction method provided in this embodiment has at least the following three major differences and technical advantages: (1) Existing models only learn and reason based on image features, which cannot cope with the problem of spatiotemporal consistency of ground feature features. The method in this embodiment proposes an explicit prior knowledge encoding method, which introduces expert experience knowledge in addition to image features into feature learning. This can strengthen the model's differential feature encoding of ground features with the same or similar image features, and improve the model's feature discrimination ability and model generalization in complex scenarios. (2) Traditional machine learning methods based on expert experience knowledge are insufficient in deep feature learning and cannot cope with complex feature spaces. The method in this embodiment proposes a prior knowledge embedding method applicable to mainstream deep learning frameworks, which enables the model to achieve integrated learning of prior knowledge, global features and local features. (3) Existing models using a single loss function training method are difficult to meet the requirements of extraction result accuracy. The method in this embodiment proposes a multi-dimensional fusion loss function, which comprehensively considers the extraction accuracy of ground features, edge details, geometric structure and uneven distribution of categories, which is conducive to the model outputting extraction results with better robustness.
[0085] To demonstrate the effectiveness of the prior embedding-based feature extraction method provided in this embodiment, some ablation experimental data are provided here for reference.
[0086] First, four cities with different latitudes and longitudes were selected as the study area. The selection criteria were: these four cities have a large geographical span (ranging from 23°N to 45°N), covering temperate to subtropical climate zones, and exhibit significant differences in architectural style, landscape features, lighting conditions (high-latitude cities have smaller incident angles and weaker sunlight, while low-latitude cities have the opposite), and atmospheric conditions. This allows for a thorough simulation of surface complexity and spatiotemporal heterogeneity, providing an ideal test scenario for evaluating model robustness. Experimental data used Sentinel-2 satellite remote sensing imagery to construct a high-resolution dataset of ground feature detection across multiple regions and time phases. Typical regional image tiles and their masks are shown below. Figure 4 As shown.
[0087] To ensure consistency in the comparison, all models were trained on the same hardware platform using a remote sensing dataset containing imagery from multiple regions and time periods. To eliminate the impact of training differences on model performance, all models used the same optimizer, loss function, and training strategy.
[0088] The experiment selected three types of region prior information for embedding learning, corresponding to Spatial Embedded (SE), Temporal Embedded (TE), and Spatial-Temporal Embedded (STE). The corresponding datasets follow the naming convention: "City ID_Imaging Timestamp_Region ID". For example, "T49QGF_20180115_HRBs_001" corresponds to city T49QGF, imaged on January 15, 2018, with region ID 001. This dataset contains four cities (T49QGF, T49SGU, T50TMK, and T51TYL), each containing labeled samples from different time points and regions.
[0089] To verify the model's generalization ability to unseen spatiotemporal scenes (new times or new regions) within the same city, the experiment employed a "spatiotemporally non-overlapping" partitioning strategy. For each city, time-region combinations were divided into training and test sets, with strict non-overlapping imaging timestamps (e.g., 20180115 and 20180311) and region IDs (e.g., 001 and 002). For example, the training set included T49QGF_20180115_001 and T49QGF_20180311_002, while the test set used untrained samples such as T49QGF_20181219_003.
[0090] To objectively validate model performance, all participating models, including the PEDNet model and its ablation variant, the U-Net model, and the FCN model, were trained and tested on the same training and test sets on this spatiotemporal dataset. The training set enabled the models to learn features of high-rise building areas (HRBs) from known urban spatiotemporal scenes, while the test set validated the models' predictive performance in unseen spatiotemporal scenes. This ensured that differences in model performance stemmed solely from structural variations, providing rigorous support for the validity of subsequent results.
[0091] The experiment consisted of two parts. First, ablation experiments on the PEDNet model focused on analyzing the impact of the SE and TE modules on the performance of the PDA attention mechanism. By enabling or disabling SE and TE, the F1 score, IoU, and OA metrics of the model were compared under different configurations. Second, comparative experiments with classic models quantitatively compared the detection results of the PEDNet, U-Net, and FCN models on the same dataset. The ablation experiment results clarified the specific roles of SE and TE in improving spatiotemporal adaptability, while the comparison with the U-Net and FCN models verified the advantages of the PEDNet model's dual-branch structure in the HRBs extraction task.
[0092] Specifically, the ablation experiments, which involved learning from the same set of training data and testing on the same data, focused on analyzing the impact of the switching states of SE and TE in the PDWA module on feature extraction. (1) PEDNet_base: Disable the SE and TE mechanisms in PDWA and retain only the basic attention structure. At this time, PDWA only relies on the original spectral and spatial features of the pixels in the window to perform self-attention calculation. It completes feature dimensionality reduction through multi-layer PDBlock (containing the PDWA basic module and MLP) and Patch Merging, without introducing any regional or temporal information for learning.
[0093] (2) PEDNet_SE: Enable the SE mechanism and disable TE in PDWA. Specifically, the SE module transforms the input regional information (such as city number) into a high-dimensional feature vector, which is then mapped to a spatial attention that matches the number of attention heads through the spatial projection layer (Spatial Proj), and participates in the calculation of attention weights within the window, enabling the model to learn the regional spatial features.
[0094] (3) PEDNet_TE: Enable the TE mechanism and disable SE in PDWA. The TE module converts the time information (year and Julian day) into a high-dimensional feature vector, and then generates time attention through the temporal projection layer (Temporal Proj), so that the time information participates in the attention modulation learning of PDWA.
[0095] (4) PEDNet_SE+TE: Enable SE and TE mechanisms simultaneously in PDWA. The two work together as dual biases for attention computation. Spatial and temporal information are learned in parallel during the window attention computation in PDWA.
[0096] All experiments were conducted on the same dataset and hardware environment. By comparing the differences in F1 score, IoU, and OA metrics among the four models, the individual and combined effects of SE and TE as built-in mechanisms of PDWA were quantified. In the ablation experiments of the PEDNet model, all configurations used the same basic parameters: window size was set to 16, the number of attention heads was 4, 8, 16, and 32 in Stage, MLP ratio was 4, and Dropout rate was 0.1. All models were trained to the same number of iterations. Four experimental models (PEDNet_base, PEDNet_SE, PEDNet_TE, and PEDNet_SE+TE) were constructed by controlling only the on / off state of SE and TE in PDWA. The model performance metrics under different configurations are shown in Table 1.
[0097] Table 1
[0098] As shown in Table 1, PEDNet_SE achieved the highest performance across all metrics, with F1, IoU, and OA values of 62.8%, 45.8%, and 91.3%, respectively. This indicates that enabling only the Spatial Encoder (SE) mechanism in PDWA is the optimal solution under the current configuration. The second best performing network is PEDNet_TE, with F1, IoU, and OA values of 61.4%, 44.3%, and 91.2%, respectively. Its outstanding OA value indicates that the model achieves good classification accuracy for the overall category when only the TE mechanism is introduced. The relatively low F1 and IoU values of PEDNet_SE+TE suggest that enabling both SE and TE mechanisms simultaneously may lead to feature interference or insufficient synergy under the current experimental settings, resulting in a decrease in the model's predictive ability for small targets and small sample categories. The worst performing network is PEDNet_base, with all metrics lower than other models that enabled embedding mechanisms, further validating the effectiveness of enabling both SE and TE in improving model performance in PDWA. In summary, the PEDNet_SE model with only the SE mechanism enabled performs better in the ablation experiments.
[0099] To further verify the model's performance in real-world scenarios, two typical samples were selected for visualization comparison: "T49SGU_20181229_HRBs_003" (approximately 34°N) was imaged in winter, and "T51TYL_20180918_HRBs_002" (approximately 45°N) was imaged in autumn. These imaging conditions, created by differences in geographical location and time, provide ideal samples for visual comparison to verify the model's robustness in complex spatiotemporal environments. The visualization results show that the PEDNet series models have significant advantages in feature extraction and detail preservation.
[0100] Figure 5The image comparison of the ablation experiment results for sample T49SGU_20181229_HRBs_003 (T49SGU region 3) is shown. From the visualization results, the PEDNet_base model's ability to capture HRBs is significantly weaker than other models incorporating embedding mechanisms, with poorer completeness and accuracy in its HRB region segmentation results. The PEDNet_SE model performs best in capturing HRBs, accurately and comprehensively identifying them, demonstrating strong feature extraction and target recognition capabilities. The PEDNet_TE and PEDNet_SE+TE models are similar in the number of HRBs captured, but the PEDNet_SE+TE model performs slightly worse than the PEDNet_TE model in detail processing, showing shortcomings in capturing some subtle HRB features. This may be due to interference from the simultaneous operation of the two embedding mechanisms, affecting the model's accuracy in extracting detailed information.
[0101] In summary, the visualization results for samples from different regions and time periods, corroborated by the tabular quantitative data, fully demonstrate that learning spatiotemporal features that integrate spatial heterogeneity and temporal differences plays a crucial role in improving the model's segmentation robustness in complex urban scenarios. Under the current experimental settings, the PEDNet_SE model with only SE enabled exhibits superior segmentation performance and overall metrics in cross-spatial and temporal scenarios.
[0102] To comprehensively verify the performance advantages of the PEDNet_SE model in semantic segmentation of remote sensing images of high-rise building areas, this invention selects the classic semantic segmentation networks U-Net and FCN as benchmarks and constructs cross-model performance evaluation experiments. U-Net, with its encoder-decoder structure and skip connection mechanism, excels in medical image segmentation; FCN, through its fully convolutional architecture, achieves end-to-end pixel-level prediction and is a fundamental model for semantic segmentation tasks. In the experiments, all models were run on the same hardware platform and trained using the same remote sensing dataset (containing multi-regional and multi-temporal images), employing the same optimizer, loss function, and training strategy to eliminate the interference of training differences on performance. By comparing the performance of each model in F1 score, IoU, and OA metrics, the ability of the PEDNet_SE model to extract categories in high-rise building areas is quantitatively analyzed. Table 2 shows the accuracy comparison of different models in the experiments.
[0103] Table 2
[0104] As shown in Table 2, the PEDNet_SE model significantly outperforms the U-Net and FCN models in F1, IoU, and OA metrics, achieving an F1 score of 62.8%, IoU of 45.8%, and OA of 91.3%, fully demonstrating its performance advantage in semantic segmentation of remote sensing images of high-rise building areas. The U-Net model has an F1 score of 54.8%, IoU of 42.3%, and OA of 90.2%; the FCN model has an F1 score of 55.8%, IoU of 38.1%, and OA of 90.1%. The comparison reveals that U-Net and FCN lag behind PEDNet_SE in all metrics, especially the IoU metric of FCN, where the difference is particularly significant.
[0105] To further verify the model's segmentation performance in real-world scenarios, and especially to highlight the advantages of the proposed model compared to classic models such as UNet and FCN in cross-temporal and cross-spatial environments, two typical samples are selected for visualization comparison: "T49SGU_20180930_HRBs_002" (approximately 34°N) is an autumn image, and "T50TMK_20180212_HRBs_002" (approximately 39°N) is a winter image. These imaging conditions, caused by latitude and seasonal differences, provide typical samples for verifying the model's segmentation robustness in cross-temporal and cross-spatial environments.
[0106] Figure 6 The image comparison of the ablation test results for sample T49SGU_20180930_HRBs_002 (T49SGU region 2) is shown. Further analysis of the visualization results reveals that the PEDNet_SE model performs better in segmenting HRBs. Within the red box area, the segmented HRBs have clear outlines, retain complete details, and accurately identify smaller building units. The U-Net model exhibits some target fragmentation and edge blurring issues in this region. The FCN model's segmentation results are relatively coarser, with insufficient ability to identify small target buildings, resulting in some high-rise buildings not being accurately segmented.
[0107] In summary, the visualization results and quantitative metrics for samples from different regions and time periods corroborate each other, fully demonstrating that the PEDNet_SE model has significant advantages over the U-Net model in terms of target fragmentation and the FCN model in terms of coarse recognition. It can more efficiently capture the contours and boundaries of small-scale building units and is more adaptable to complex scenes. This advantage is reflected not only in the superior quantitative metrics but also in the visual results, showing that under the current experimental settings, the PEDNet_SE model outperforms classic models in segmentation performance and various metrics across spatiotemporal scenes.
[0108] This invention also provides a device for extracting ground features based on prior embedding. Please refer to [link / reference]. Figure 7 , Figure 7 This is a schematic diagram of the prior embedding-based feature extraction device provided by the present invention. In this embodiment, the prior embedding-based feature extraction device includes an acquisition module 710 and a feature extraction module 720.
[0109] The acquisition module 710 is used to acquire optical satellite imagery and regional prior information of optical satellite imagery.
[0110] Optical satellite imagery is obtained by taking pictures of a target area, and the prior information of the area is information that characterizes the features of the ground features in the target area.
[0111] The feature extraction module 720 is used to input optical satellite imagery and regional prior information into the feature extraction model to obtain the feature extraction results of the optical satellite imagery output by the feature extraction model.
[0112] The feature extraction model is trained based on sample optical satellite imagery, prior information of the sample area corresponding to the sample optical satellite imagery, and the feature extraction results of the sample optical satellite imagery.
[0113] In some embodiments, the ground feature extraction model includes an attention feature extraction network, a convolutional feature extraction network, a feature pyramid network, and a feature fusion network. The attention feature extraction network, the feature pyramid network, and the feature fusion network are connected sequentially, and the convolutional feature extraction network and the feature fusion network are connected. The convolutional feature extraction network is used to extract spatial features and local semantic features from optical satellite imagery to generate local image features. The attention feature extraction network and the feature pyramid network are used to embed prior information and extract global semantic features from optical satellite imagery and regional prior information to generate global image features. The feature fusion network is used to generate ground feature extraction results based on local and global image features.
[0114] In some embodiments, the attention feature extraction network includes a Stem module, a first stage module, a second stage module, a third stage module, and a fourth stage module connected in sequence, and the feature pyramid network includes a first network layer, a second network layer, a third network layer, and a fourth network layer; wherein, the first stage module, the second stage module, the third stage module, and the fourth stage module all include a PEBlock module, and the PEBlock module includes a prior embedding window attention module; the first stage module is connected to the first network layer, the second stage module is connected to the second network layer, the third stage module is connected to the third network layer, and the fourth stage module is connected to the fourth network layer.
[0115] In some embodiments, the global image features include a first global feature map, a second global feature map, a third global feature map, and a fourth global feature map; wherein, the Stem module is used to extract features from the optical satellite image to generate an initial feature map; the first Stage module is used to embed prior information and extract global semantic features from the initial feature map and regional prior information to generate a first feature map; a first network layer is used to perform convolution and upsampling processing on the first feature map to generate a first global feature map; and the second Stage module is used to embed prior information and extract global semantic features from the first feature map and regional prior information. The process involves generating a second feature map; a second network layer performing convolution and upsampling on the second feature map to generate a second global feature map; a third stage module embedding prior information and extracting global semantic features from the second feature map and region prior information to generate a third feature map; a third network layer performing convolution and upsampling on the third feature map to generate a third global feature map; a fourth stage module embedding prior information and extracting global semantic features from the third feature map and region prior information to generate a fourth feature map; and a fourth network layer performing convolution on the fourth feature map to generate a fourth global feature map.
[0116] In some embodiments, the feature fusion network includes a fifth network layer and a sixth network layer; wherein, the fifth network layer is used to perform feature fusion on the first global feature map, the second global feature map, the third global feature map, the fourth global feature map and local image features to generate a fused feature map; the sixth network layer is used to perform upsampling processing on the fused feature map to generate ground feature extraction results.
[0117] In some embodiments, the feature extraction model is trained based on a preset loss function; wherein the preset loss function is determined based on geometric constraint loss and edge loss, and the edge loss is determined based on cross-entropy dice joint loss, binary cross-entropy loss and focal loss.
[0118] In some embodiments, the feature extraction model is trained based on the following steps: acquiring sample optical satellite images, prior information of the sample area corresponding to the sample optical satellite images, and the feature extraction results of the sample optical satellite images; and training the initial PEDNet model based on the sample optical satellite images, prior information of the sample area, and the feature extraction results of the sample optical satellite images to obtain the feature extraction model.
[0119] The present invention also provides an electronic device. Figure 8 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 8As shown, the electronic device may include a processor 810, a communications interface 820, a memory 830, and a communication bus 840. The processor 810, communications interface 820, and memory 830 communicate with each other via the communication bus 840. The processor 810 can call logical instructions from the memory 830 to execute a method for extracting ground features based on prior embedding.
[0120] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0121] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the prior embedding-based method for extracting ground features provided by the above methods.
[0122] The present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the prior embedding-based ground feature extraction method provided by the above methods.
[0123] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0124] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0125] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for extracting ground features based on prior embedding, characterized in that, include: Acquire optical satellite imagery and prior regional information of the optical satellite imagery; The optical satellite imagery is obtained by taking pictures of the target area, and the prior information of the area is information that characterizes the features of the ground features in the target area. The optical satellite imagery and the prior information of the region are input into the ground feature extraction model to obtain the ground feature extraction results of the optical satellite imagery output by the ground feature extraction model. The ground feature extraction model is trained based on sample optical satellite imagery, prior information of the sample area corresponding to the sample optical satellite imagery, and the sample ground feature extraction results corresponding to the sample optical satellite imagery.
2. The method for extracting ground features based on prior embedding according to claim 1, characterized in that, The ground feature extraction model includes an attention feature extraction network, a convolutional feature extraction network, a feature pyramid network, and a feature fusion network. The attention feature extraction network, the feature pyramid network, and the feature fusion network are connected in sequence, and the convolutional feature extraction network and the feature fusion network are connected. The convolutional feature extraction network is used to extract spatial features and local semantic features from the optical satellite imagery to generate local image features. The attention feature extraction network and the feature pyramid network are used to embed prior information and extract global semantic features from the optical satellite image and the regional prior information to generate global image features. The feature fusion network is used to generate the extracted ground feature results based on the local image features and the global image features.
3. The method for extracting ground features based on prior embedding according to claim 2, characterized in that, The attention feature extraction network includes a Stem module, a first stage module, a second stage module, a third stage module, and a fourth stage module connected in sequence, and the feature pyramid network includes a first network layer, a second network layer, a third network layer, and a fourth network layer. The first Stage module, the second Stage module, the third Stage module, and the fourth Stage module all include a PEBlock module, and the PEBlock module includes a priori embedded window attention module. The first Stage module is connected to the first network layer, the second Stage module is connected to the second network layer, the third Stage module is connected to the third network layer, and the fourth Stage module is connected to the fourth network layer.
4. The method for extracting ground features based on prior embedding according to claim 3, characterized in that, The global image features include a first global feature map, a second global feature map, a third global feature map, and a fourth global feature map; The Stem module is used to extract features from the optical satellite imagery and generate an initial feature map. The first Stage module is used to embed prior information and extract global semantic features from the initial feature map and the region prior information to generate a first feature map; The first network layer is used to perform convolution and upsampling on the first feature map to generate the first global feature map; The second Stage module is used to embed prior information and extract global semantic features from the first feature map and the region prior information to generate a second feature map; The second network layer is used to perform convolution and upsampling on the second feature map to generate the second global feature map; The third Stage module is used to embed prior information and extract global semantic features from the second feature map and the region prior information to generate the third feature map. The third network layer is used to perform convolution and upsampling on the third feature map to generate the third global feature map; The fourth Stage module is used to embed prior information and extract global semantic features from the third feature map and the region prior information to generate the fourth feature map. The fourth network layer is used to perform convolution processing on the fourth feature map to generate the fourth global feature map.
5. The method for extracting ground features based on prior embedding according to claim 4, characterized in that, The feature fusion network includes a fifth network layer and a sixth network layer; The fifth network layer is used to perform feature fusion on the first global feature map, the second global feature map, the third global feature map, the fourth global feature map and the local image features to generate a fused feature map. The sixth network layer is used to upsample the fused feature map to generate the extracted ground feature results.
6. The method for extracting ground features based on prior embedding according to claim 1, characterized in that, The feature extraction model is trained based on a preset loss function. The preset loss function is determined based on geometric constraint loss and edge loss, and the edge loss is determined based on cross-entropy dice joint loss, binary cross-entropy loss and focus loss.
7. The method for extracting ground features based on prior embedding according to claim 1, characterized in that, The feature extraction model is trained based on the following steps: Acquire sample optical satellite images, prior information of the sample area corresponding to the sample optical satellite images, and extraction results of sample ground features corresponding to the sample optical satellite images; Based on the sample optical satellite imagery, the prior information of the sample area, and the extraction results of the sample ground features, the initial PEDNet model is trained to obtain the ground feature extraction model.
8. A device for extracting ground features based on prior embedding, characterized in that, include: The acquisition module is used to acquire optical satellite imagery and the region prior information of the optical satellite imagery; The optical satellite imagery is obtained by taking pictures of the target area, and the prior information of the area is information that characterizes the features of the ground features in the target area. The feature extraction module is used to input the optical satellite image and the prior information of the region into the feature extraction model to obtain the feature extraction result of the optical satellite image output by the feature extraction model; The ground feature extraction model is trained based on sample optical satellite imagery, prior information of the sample area corresponding to the sample optical satellite imagery, and the sample ground feature extraction results corresponding to the sample optical satellite imagery.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the method for extracting ground features based on prior embedding as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the prior embedding-based feature extraction method as described in any one of claims 1 to 7.