Regional earth surface element remote sensing extraction-oriented domain migration method, device, equipment and medium

By using multiple reconstruction sample images and teacher network models from different local spatial perspectives in the domain migration method, the problem of poor optimization effect of traditional domain migration method is solved, and more accurate remote sensing extraction of geographic elements is achieved.

CN119964163AActive Publication Date: 2025-05-09AEROSPACE INFORMATION RES INST CAS

Patent Information

Application Number
CN202510437532.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-05-09
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

The traditional domain migration method in the prior art does not have the best optimization effect on the trained model, which makes it difficult to accurately remote sensing the geometric elements extract the test data when there are differences in feature distribution between the training sample data and the test data.

Method used

By obtaining the remote sensing image of the target area, multiple reconstructed sample images with different local spatial perspectives were obtained, and the teacher network model was used to perform improved domain migration training on the student network model, and the trained student network model was obtained as the remote sensing extraction model of the land feature.

Benefits of technology

The adaptability of the source domain model to the target area has been significantly improved, and the accuracy of remote sensing extraction of land objects in the target area has been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964163A_ABST
    Figure CN119964163A_ABST
Patent Text Reader

Abstract

The invention provides a domain migration method, device and equipment for remote sensing extraction of regional earth surface elements and a medium, and relates to the technical field of image processing, and the method comprises the steps: respectively determining source domain models as a teacher network model and a student network model, and based on each reconstructed sample image and the teacher network model, obtaining a source domain model; carrying out improved domain migration training on the student network model to obtain a trained student network model, and taking the trained student network model as a ground feature element remote sensing extraction model corresponding to the target region; and obtaining a ground feature element remote sensing extraction result of the target area based on the ground feature element remote sensing extraction model corresponding to the target area and the remote sensing image of the target area. According to the method, domain migration can be carried out on the source domain model through improved domain migration training, the features in the remote sensing image of the target area are migrated to the source domain model, the adaptive capacity of the source domain model to the target area can be remarkably improved, and the accuracy of ground feature element remote sensing extraction can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a domain migration method, device, equipment and medium for remote sensing extraction of regional surface elements. Background Art

[0002] Remote sensing extraction of land features refers to the process of using remote sensing images to obtain information such as the number and spatial distribution of land features such as buildings, roads, water bodies, cultivated land and forest land on the earth's surface. Remote sensing extraction of land features has been widely used in many fields such as resource management, environmental testing, urban planning and disaster monitoring.

[0003] The traditional domain transfer method in the related art can use the model trained based on the training sample data and sample labels to realize the remote sensing extraction of ground features of the test data. However, due to the influence of spatiotemporal heterogeneity, when the training sample data and the test data are respectively obtained based on remote sensing data corresponding to different regions, different times or different surface morphologies, there are differences in feature distribution between the training sample data and the test data, which do not obey the independent and identically distributed assumption of machine learning, resulting in low accuracy of remote sensing extraction of ground features of the test data based on the above model, and problems such as insufficient model generalization.

[0004] Domain migration technology can migrate the knowledge of training sample data to test data in scenarios where there are differences in feature distribution between training sample data and test data (domain shift) and training sample data is missing, relying only on models trained with training sample data and unlabeled test data, so as to improve the generalization ability of the model in test data without adding additional annotation costs. However, the optimization effect of optimizing the trained model based on traditional domain migration technology in related technologies is not good. Based on the optimized model, it is still difficult to accurately perform remote sensing extraction of ground object elements on test data when there are differences in feature distribution between training sample data and test data. Therefore, how to better optimize the trained model, so as to more accurately perform remote sensing extraction of ground object elements on test data when there are differences in feature distribution between training sample data and test data, is a technical problem that needs to be solved urgently in this field. Summary of the invention

[0005] The present invention provides a domain migration method, device, equipment and medium for remote sensing extraction of regional surface elements, which is used to solve the defects that the traditional domain migration technology in the prior art has a poor optimization effect on the trained model, and it is still difficult to accurately perform remote sensing extraction of land object elements on the test data based on the optimized model when there is a difference in feature distribution between the training sample data and the test data, so as to achieve better optimization of the trained model, thereby more accurately performing remote sensing extraction of land object elements on the test data when there is a difference in feature distribution between the training sample data and the test data.

[0006] The present invention provides a domain migration method for remote sensing extraction of regional surface elements, comprising the following steps.

[0007] Acquire a remote sensing image of a target area as a target domain image; based on the target domain image, obtain a plurality of reconstructed sample images with different local spatial perspectives of the target area; based on each of the reconstructed sample images and the teacher network model, perform improved domain transfer training on the student network model to obtain a trained student network model as a remote sensing extraction model of land features corresponding to the target area, the model structures of the teacher network model and the student network model are the same as the model structures of the source domain model, the initial model parameters of the teacher network model and the student network model are the model parameters of the source domain model, the source domain model is obtained by training based on the remote sensing image of the sample area and the remote sensing image of the sample area after annotation, the annotation information in the remote sensing image of the sample area after annotation is used to indicate at least one of the type, identification, quantity and area range of the land features; based on the remote sensing extraction model of land features corresponding to the target area and the remote sensing image of the target area, obtain the remote sensing extraction result of land features of the target area.

[0008] According to a domain migration method for remote sensing extraction of regional surface elements provided by the present invention, based on the target domain image, a plurality of reconstructed sample images with different local spatial perspectives of the target area are obtained, including: based on a sliding window of a preset size and a predefined overlap rate, the target domain image is cropped to obtain a plurality of original sample images; for each original sample image, a plurality of sub-region boundary boxes of the same size, located in different regions of each original sample image and overlapping with each other are randomly generated in each original sample image; based on each of the sub-region boundary boxes in each original sample image and the boundary box of each original sample image, the region of interest alignment ROIAlign technology is used to obtain each reconstructed sample image corresponding to each original sample image.

[0009] According to a domain migration method for regional surface element remote sensing extraction provided by the present invention, based on each of the reconstructed sample images and the teacher network model, an improved domain migration training is performed on the student network model to obtain a trained student network model as a remote sensing extraction model of the surface element corresponding to the target area, including: In the k In the training, based on the reconstructed sample images, the k The teacher network model and the student network model in the training are obtained. k Each pseudo label corresponding to each original sample image in the training kThe embedded features corresponding to each reconstructed sample image corresponding to each original sample image in the training and the embedded features corresponding to each original sample image; Based on k Each pseudo label corresponding to each original sample image in the training is constructed k The spatiotemporal consistency loss function corresponding to the training is based on the The embedded features corresponding to each reconstructed sample image corresponding to each original sample image in the training and the embedded features corresponding to each original sample image are constructed to obtain the first k The spatiotemporal contrast loss function corresponding to the training; Based on The spatiotemporal consistency loss function and the spatiotemporal contrast loss function corresponding to the training are combined with the minimum entropy loss function to construct the k The target loss function corresponding to the training; Calculate the The function value of the target loss function corresponding to the training is based on the k The function value of the target loss function corresponding to the training time determines whether the student network model has not converged or k When the number of training times is not greater than the maximum, the model parameters of the student network model are updated. k Increase by 1 and return to execute the k In the training, based on the reconstructed sample images, the k The teacher network model and the student network model in the training are obtained. k Each pseudo label corresponding to each original sample image in the training k The steps of embedding features corresponding to each reconstructed sample image corresponding to each original sample image and embedding features corresponding to each original sample image in the first training are based on the first k The function value of the target loss function corresponding to the training time determines whether the student network model converges or k When the number of training times is greater than the maximum number, it is determined that the student network model has been trained, and the trained student network model is determined as the remote sensing extraction model of the land feature corresponding to the target area.

[0010] According to a domain migration method for remote sensing extraction of regional surface elements provided by the present invention, based on each of the reconstructed sample images, using the k The teacher network model and the student network model in the training are obtained. k Each pseudo label corresponding to each original sample image in the training includes: In the k In the training, the model parameters of the teacher network model are updated based on the exponential moving average method; The reconstructed sample images corresponding to each original sample image are input into the teacher network model and the student network model respectively, and the first k The first predicted image of each reconstructed sample image corresponding to each original sample image in the training and the first predicted image output by the student network model k A second predicted image of each reconstructed sample image corresponding to each original sample image in the training; Using the reverse process of ROI Align technology, k The first predicted image of each reconstructed sample image corresponding to each original sample image in the training is fused to obtain the first k The fused prediction image corresponding to each original sample image in the training; Based on k The fused prediction image corresponding to each original sample image in the training is obtained by using ROI Align technology. k Each secondary reconstructed image corresponding to each original sample image in the training is used as the k Each pseudo label corresponding to each original sample image in the training.

[0011] According to a domain migration method for remote sensing extraction of regional surface elements provided by the present invention, the domain migration method based on the first k Each pseudo label corresponding to each original sample image in the training is constructed k The spatiotemporal consistency loss function corresponding to the training includes: In the k In the training, each original sample image is input into the final target history teacher network model to obtain the output of the final target history teacher network model. k The third predicted image corresponding to each original sample image in the training; Based on k The third predicted image and the first k The fused prediction image corresponding to each original sample image in the training is calculated to obtain the k A consistency weight map corresponding to each original sample image in the training; Based on k The consistency weight map corresponding to each original sample image in the training and k Each pseudo label corresponding to each original sample image in the training is constructed to obtain the k The spatiotemporal consistency loss function corresponding to the training; The last target history teacher network model is the last target history teacher network model in the history teacher network model queue; The history teacher network model queue is obtained based on the following steps: When it is determined that the first training is finished, the teacher network model in the first training is determined as the target history teacher network model corresponding to the first training, and the target history teacher network model corresponding to the first training is added to the history teacher network model queue; In determining the m The training is completed and confirmed m When the difference between the number of training times corresponding to the target history teacher network model ranked first in the history teacher network model queue is a preset value, the first m The teacher network model in the training is determined as the target historical teacher network model. ; When the number of target history teacher networks in the history teacher network model queue is less than the number threshold, m The target history teacher network model corresponding to the training is inserted into the first position of the history teacher network model queue. When the number of target history teacher networks in the history teacher network model queue is equal to the number threshold, the first m The target history teacher network model corresponding to the training is inserted into the first position of the history teacher network model queue, and the target history teacher network model at the last position in the history teacher network model queue is removed.

[0012] According to a domain migration method for remote sensing extraction of regional surface elements provided by the present invention, based on each of the reconstructed sample images, using the k The teacher network model and the student network model in the training are obtained. k Each pseudo label corresponding to each original sample image in the training k The embedded features corresponding to each reconstructed sample image corresponding to each original sample image in the training and the embedded features corresponding to each original sample image include: In the k In the training, after inputting each reconstructed sample image corresponding to each original sample image into the student network model, the output of the feature extractor in the student network model is obtained. In the training, each original sample image corresponds to each reconstructed sample image corresponding to the feature image, k In the training, after each original sample image is input into the final target history teacher network model, the feature extractor output of the final target history teacher network model is obtained. kThe feature image corresponding to each original sample image in the training; The first k The feature image corresponding to each reconstructed sample image corresponding to each original sample image and the feature image corresponding to each original sample image in the training are projected into a feature space of a preset dimension to obtain the first k The embedded features corresponding to each reconstructed sample image corresponding to each original sample image in the training and the embedded features corresponding to each original sample image.

[0013] The present invention also provides a domain migration device for remote sensing extraction of regional surface elements, including the following modules.

[0014] A data acquisition module is used to acquire a remote sensing image of a target area as a target domain image; a sample reconstruction module is used to perform improved domain transfer training on a student network model based on each of the reconstructed sample images and the teacher network model to obtain a trained student network model as a remote sensing extraction model of land features corresponding to the target area, wherein the model structures of the teacher network model and the student network model are the same as the model structures of the source domain model, and the initial model parameters of the teacher network model and the student network model are the model parameters of the source domain model, and the source domain model is obtained by training based on the remote sensing image of the sample area and the remote sensing image of the labeled sample area, and the annotation information in the remote sensing image of the labeled sample area is used to indicate at least one of the type, identification, quantity and area range of the land features; a remote sensing extraction module is used to acquire the remote sensing extraction result of the land features of the target area based on the remote sensing extraction model of the land features corresponding to the target area and the remote sensing image of the target area.

[0015] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the domain migration method for remote sensing extraction of regional surface elements as described in any one of the above methods is implemented.

[0016] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the domain migration method for remote sensing extraction of regional surface elements as described in any one of the above is implemented.

[0017] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements any of the above-mentioned domain migration methods for remote sensing extraction of regional surface elements.

[0018] The domain migration method, device, equipment and medium for remote sensing extraction of regional surface elements provided by the present invention obtain a remote sensing image of the target area as the target domain image, and then obtain multiple reconstructed sample images with different local spatial perspectives of the target area based on the target domain image. Based on each reconstructed sample image and the teacher network model, an improved domain migration training is performed on the student network model to obtain the trained student network model as the remote sensing extraction model of the object elements corresponding to the target area, and then based on the remote sensing extraction model of the object elements corresponding to the target area and the remote sensing image of the target area, the remote sensing extraction result of the object elements of the target area is obtained. In the case where the training sample data used to train the source domain model is missing, the source domain model can be domain migrated through the improved domain migration training, and the features in the remote sensing image of the target area can be migrated to the source domain model, which can significantly improve the adaptability of the source domain model to the target area, thereby improving the accuracy of remote sensing extraction of object elements in the target area. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0020] Figure 1 This is one of the flow charts of the domain migration method for remote sensing extraction of regional surface elements provided by the present invention.

[0021] Figure 2 This is the second flow chart of the domain migration method for remote sensing extraction of regional surface elements provided by the present invention.

[0022] Figure 3 This is the third flow chart of the domain migration method for remote sensing extraction of regional surface elements provided by the present invention.

[0023] Figure 4 It is a framework diagram for constructing a spatiotemporal consistency loss function in the domain migration method for remote sensing extraction of regional surface elements provided by the present invention.

[0024] Figure 5 This is one of the effect comparison diagrams of the domain migration method for regional surface element remote sensing extraction provided by the present invention and the traditional domain migration method for regional surface element remote sensing extraction in the related art.

[0025] Figure 6 This is the second comparison diagram of the effects of the domain migration method for remote sensing extraction of regional surface elements provided by the present invention and the traditional domain migration method for remote sensing extraction of regional surface elements in the related art.

[0026] Figure 7 It is a structural schematic diagram of a domain migration device for remote sensing extraction of regional surface elements provided by the present invention.

[0027] Figure 8 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0028] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0029] In the description of the invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0030] In the description of the present application, the terms "first", "second", etc. are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are usually a class, and the number of objects is not limited. For example, the first object can be one or more. In addition, in the description of the present application, "and / or" represents at least one of the connected objects, and the character " / " generally represents that the front and back associated objects are in an "or" relationship.

[0031] It should be noted that the methods for optimizing the trained ground feature extraction model based on domain transfer technology mainly include the following three categories: domain transfer methods based on image transfer, domain transfer methods based on adversarial training, and domain transfer methods based on self-training. Among them, the domain transfer method based on self-training is the most commonly used method in practical applications.

[0032] The domain transfer method based on self-training generally uses the trained model to generate pseudo labels for test data, and uses the pseudo labels to iteratively fine-tune the model, thereby gradually adapting the knowledge of the training sample data to the test data. Therefore, the quality of the pseudo labels is the key to affecting the adaptability of the model.

[0033] However, in the related art, when optimizing the model based on the self-training domain transfer method, the limitation of defining the reliability of pseudo labels only based on the output information of the current model iteration is not considered. Among them, the unreliability of pseudo labels mainly manifests in two aspects, namely, pseudo label inconsistency and knowledge forgetting.

[0034] On the one hand, experimental observations show that the model is prone to produce inconsistent semantic categories under different local spatial perspectives, and the semantic inconsistency generated by pseudo-labels during the iteration process is a key factor leading to unstable model training and even performance degradation.

[0035] On the other hand, knowledge forgetting is another key factor that causes instability in the model optimization process. Since the self-training process lacks effective labeled data, model knowledge becomes the only available supervisory information for the domain transfer process. As the training process continues to iterate, the model will gradually deviate from the knowledge in the training sample data, resulting in the output information of the current model iteration being inaccurate.

[0036] Therefore, the optimization effect of optimizing the trained model based on the traditional domain transfer technology in the related art is not good. Based on the optimized model, it is still difficult to accurately perform remote sensing extraction of ground object elements on the test data when there is a difference in feature distribution between the training sample data and the test data. How to better optimize the trained model so as to more accurately perform remote sensing extraction of ground object elements on the test data when there is a difference in feature distribution between the training sample data and the test data is a technical problem to be solved in this field.

[0037] In this regard, the present invention provides a domain migration method for remote sensing extraction of land features in a large area. The land feature remote sensing method provided by the present invention proposes a spatial multi-perspective consistency learning mechanism to address the prediction inconsistency problem of the source domain model under different spatial contexts, and obtains a stable pseudo-label unified from multiple perspectives through a spatial multi-perspective enhancement and prediction fusion process; in order to avoid the problem of knowledge forgetting caused by the lack of supervisory information guidance in the domain migration process, a temporal dynamic consistency learning mechanism is proposed, and the model knowledge of historical moments is introduced to regularize the training direction of the model, so that the model has the ability to self-correct the current pseudo-label noise; at the same time, in order to alleviate the cognitive bias problem caused by domain shift in the large-area surface feature extraction and classification task, a spatiotemporal comparison strategy is designed, and contrast learning is used to constrain the intrinsic semantic association of different category features, so as to improve the category discrimination representation of the surface feature remote sensing extraction and classification model in the target domain.

[0038] Combine the following Figure 1-Figure 6 The present invention describes a domain migration method for remote sensing extraction of regional surface elements.

[0039] Figure 1 This is one of the flow charts of the domain migration method for remote sensing extraction of regional surface elements provided by the present invention, such as Figure 1 As shown, the method includes the following steps: Step 101, obtaining a remote sensing image of a target area as a target domain image.

[0040] It should be noted that the execution subject of the embodiment of the present invention is a domain migration device for remote sensing extraction of regional surface elements. The domain migration device for remote sensing extraction of regional surface elements can be configured in electronic devices such as computers or servers.

[0041] Specifically, the target area is the remote sensing extraction object of the domain migration method for remote sensing extraction of regional surface elements provided by the present invention. Based on the domain migration method for remote sensing extraction of regional surface elements provided by the present invention, the type, quantity and spatial distribution of the ground features in the target area can be obtained based on the remote sensing image of the target area.

[0042] It should be noted that the geographical features in the embodiments of the present invention may include but are not limited to water systems, settlements and facilities, transportation facilities, pipeline facilities, boundaries, landforms, vegetation, and soil quality. Water systems may include but are not limited to natural water bodies such as rivers, lakes, and reservoirs, and their ancillary facilities, such as dams and bridges. Settlements and facilities may include but are not limited to cities, villages, buildings, and public facilities. Transportation facilities may include but are not limited to transportation facilities such as roads, railways, bridges, and tunnels, and their ancillary facilities. Pipeline facilities may include but are not limited to various pipelines and lines, such as oil pipelines and power transmission lines. Boundaries may include but are not limited to administrative boundaries such as national boundaries, provincial boundaries, and county boundaries. Vegetation and soil quality may include but are not limited to vegetation-covered areas such as forests and grasslands, as well as different types of soil.

[0043] It is understandable that the target area in the embodiment of the present invention may be determined based on actual needs. The target area in the embodiment of the present invention is not specifically limited.

[0044] It should be noted that the target area in the embodiment of the present invention may be an area with a relatively large area, for example, the target area may be an area under the jurisdiction of a certain administrative unit.

[0045] It should be noted that the remote sensing image of the target area in the embodiment of the present invention may be collected by a remote sensing satellite or by an image sensor disposed on a drone. The specific method for obtaining the remote sensing image of the target area in the embodiment of the present invention is not limited.

[0046] In the embodiment of the present invention, the remote sensing image of the target area can be obtained in a variety of ways, for example, the remote sensing image of the target area can be obtained based on the user's input; or the remote sensing image of the target area sent by other electronic devices can be received. The specific method of obtaining the remote sensing image of the target area is not limited in the embodiment of the present invention.

[0047] After acquiring the remote sensing image of the target area, the remote sensing image of the target area can be determined as the target domain image data.

[0048] Step 102: Based on the target domain image, a plurality of reconstructed sample images with different local spatial perspectives of the target area are obtained.

[0049] Specifically, after acquiring the target domain image, multiple reconstructed sample images with different local spatial perspectives can be obtained based on the target domain image through image preprocessing and deep learning technology.

[0050] It should be noted that the different local spatial perspectives in the embodiments of the present invention can be understood as different spatial positions, observation angles, and observation scales when observing a part of the target area. Accordingly, the multiple reconstructed sample images with different local spatial perspectives may include images of multiple different local areas in the target area, and at least one of the shooting positions, shooting angles, and shooting scales of the images of the local areas is different.

[0051] As an optional embodiment, based on the target domain image, multiple reconstructed sample images with different local spatial perspectives of the target area are obtained, including: based on a sliding window of a preset size and a predefined overlap rate, the target domain image is cropped to obtain multiple original sample images.

[0052] Specifically, the sliding window in the embodiment of the present invention is a rectangle, and the length of the sliding window is and width is preset based on prior knowledge and / or actual conditions. In the embodiment of the present invention, the size of the sliding window ( ) are not specifically limited.

[0053] Overlap rate in the embodiment of the present invention Refers to the overlap ratio of the above sliding windows before and after sliding. The specific value of can be predefined based on prior knowledge and / or actual conditions. There is no restriction on the specific value of .

[0054] Based on the above preset size and predefined overlap rate, the sliding step length of the above sliding window can be calculated, and then starting from the upper left corner of the target domain image, the above sliding window can be slid according to the above sliding step length, and the position of the above sliding window after each sliding is determined as the position of each original sample image in the target domain image.

[0055] After determining the position of each original sample image in the target domain image, the target domain image can be cropped to obtain each original sample image. Each original sample image can be recorded as .in, Indicates The original sample image, Indicates not greater than A positive integer of ; Represents the total number of original sample images; Indicates Original sample images The size is , the number of channels is 3.

[0056] For each original sample image, a plurality of sub-region boundary boxes with the same size, located in different regions of each original sample image and overlapping with each other are randomly generated in each original sample image.

[0057] Specifically, each original sample image can be recorded as The Original sample images , can be found in Original sample images Randomly generated The same size, located in Original sample images The bounding boxes of the sub-regions in different regions that overlap with each other. represents a positive integer greater than 2, The value of can be determined based on prior knowledge and / or actual conditions, for example The value of is 4. There is no restriction on the specific value of .

[0058] No. Original sample images The The sub-region bounding box can be expressed as , Indicates not greater than A positive integer; Original sample images The bounding boxes of each sub-region in can be expressed as .

[0059] Based on the bounding boxes of each sub-region in each original sample image and the bounding box of each original sample image, the reconstructed sample images corresponding to each original sample image are obtained by using the region of interest alignment ROI Align technology.

[0060] It should be noted that ROI Align technology is a technology widely used in the field of target detection and image recognition. It is mainly used to solve the position mismatch problem caused by quantization operation in the region of interest pooling (ROI Pooling). The core idea of ​​ROI Align technology is to cancel the quantization operation and use the bilinear interpolation method to obtain the image value at the pixel point with floating point coordinates, thereby converting the entire feature aggregation process into a continuous operation.

[0061] The first Original sample images middle The sub-region bounding boxes and sizes are No. Original sample images After the bounding box of is input into the ROI Align network, the output of the above ROI Align network can be obtained. Original sample images The corresponding reconstructed sample images. Original sample images The size of each corresponding reconstructed sample image is , No. Original sample images The corresponding number of reconstructed sample images is open.

[0062] No. Original sample images The corresponding The reconstructed sample image can be expressed as , Indicates not greater than A positive integer. Original sample images The corresponding reconstructed sample images can be expressed as .

[0063] Step 103: Based on the reconstructed sample images and the teacher network model, the student network model is improved in domain transfer training to obtain a trained student network model as the remote sensing extraction model of the land features corresponding to the target area. The model structure of the teacher network model and the student network model is the same as the model structure of the source domain model. The initial model parameters of the teacher network model and the student network model are the model parameters of the source domain model. The source domain model is trained based on the remote sensing images of the sample area and the remote sensing images of the labeled sample area. The annotation information in the labeled remote sensing images of the sample area is used to indicate at least one of the type, identification, quantity and area range of the land features.

[0064] Specifically, after obtaining the remote sensing image of the sample area and the remote sensing image of the labeled sample area, the size of Sliding window and overlap ratio , calculate the sliding step length of the sliding window, and then start from the upper left corner of the remote sensing image of the sample area and the remote sensing image of the labeled sample area respectively, slide the sliding window according to the sliding step length, and determine the position of the sliding window after each sliding as the position of each source domain sample image in the remote sensing image of the sample area and the position of each labeled source domain sample image in the remote sensing image of the labeled sample area.

[0065] After determining the position of each source domain sample image in the remote sensing image of the sample area and the position of each labeled source domain sample image in the remote sensing image of the labeled sample area, the remote sensing image of the sample area and the labeled sample area can be cropped to obtain each source domain sample image. Each source domain sample image and each labeled source domain sample image can be recorded as .in, Indicates Zhang source domain sample image, Indicates the first Zhang source domain sample image, Indicates not greater than A positive integer of ; Represents the total number of sample images in each source domain; Indicates Sample image of source domain The size is , the number of channels is 3; Indicates the first Sample image of source domain The size is , the number of channels is 3.

[0066] The source domain sample images are used as training sample data, and the annotated source domain sample images are used as sample labels to perform supervised training on the initial model. The training loss is the pixel-level cross entropy loss. , we can obtain the trained source domain model .in, Represents the model parameters of the trained source domain model.

[0067] It should be noted that the initial model in the embodiment of the present invention can be constructed based on a traditional deep learning network model. For example, the initial model can be constructed based on a SegFormer model (a deep learning model based on a Transformer architecture).

[0068] The initial model may include a feature extractor and a classifier. The feature extractor may be used to extract the features of the input remote sensing image to obtain a feature image corresponding to the remote sensing image. The classifier may be used to identify and classify the ground object elements based on the feature image, thereby outputting the remote sensing recognition result of the ground object elements of the input image.

[0069] It is understandable that in the embodiment of the present invention, the sample area may be the same as or different from the target area, but the remote sensing image of the sample area and the remote sensing image of the target area are remote sensing images of different phases.

[0070] In the embodiment of the present invention, the trained source domain model can be , respectively determined as the teacher network model and student network model , the model parameters of the trained source domain model Determined as the teacher network model The initial values ​​of the model parameters and student network model Model parameters Initial value of .

[0071] Figure 2 This is the second flow chart of the domain migration method for remote sensing extraction of regional surface elements provided by the present invention. Figure 2 As shown, as an optional embodiment, based on each reconstructed sample image and the teacher network model, the student network model is improved in domain transfer training to obtain a trained student network model as a remote sensing extraction model of ground features corresponding to the target area, including: k In the training, based on each reconstructed sample image, using the k The teacher network model and student network model in the training are obtained. k Each pseudo label corresponding to each original sample image in the training.

[0072] Based on each reconstructed sample image, using the k The teacher network model and student network model in the training are obtained. k Each pseudo label corresponding to each original sample image in the training k The embedded features corresponding to each reconstructed sample image and the embedded features corresponding to each original sample image in the training are obtained.

[0073] Based on k Each pseudo label corresponding to each original sample image in the training is constructed k The spatiotemporal consistency loss function corresponding to the training is based on the k The embedding features corresponding to each reconstructed sample image and the embedding features corresponding to each original sample image in the training are constructed to obtain the first k The spatiotemporal contrast loss function corresponding to the training.

[0074] Figure 3 This is the third flow chart of the domain migration method for remote sensing extraction of regional surface elements provided by the present invention. Figure 4 This is a framework diagram for constructing a spatiotemporal consistency loss function in the domain migration method for regional surface element remote sensing extraction provided by the present invention. Figure 3 and Figure 4 As shown, as an optional embodiment, based on each reconstructed sample image, using the k The teacher network model and student network model in the training are obtained. k Each pseudo label corresponding to each original sample image in the training includes: k In this training, the model parameters of the teacher network model are updated based on the exponential moving average method.

[0075] It should be noted that the student network model in the embodiment of the present invention is a trainable network, and the teacher network model uses a dynamic update mechanism of a dynamic teacher-student teacher (DT-ST) to update the model parameters. The specific update method uses an exponential moving average (EMA) method. The mathematical expression of the model parameter update of the teacher network model is as follows:

[0076] in, Represents the smoothing coefficient (usually close to 1, such as 0.99), which is used to control the update speed; Indicates k The model parameters of the teacher network model in the training; Indicatesk The model parameters of the student network model in the training; Indicates km The model parameters of the teacher network model in the training; m Represents a predefined integer value, such as m The value can be 3.

[0077] It should be noted that in In the case of km The value of k .

[0078] Input each reconstructed sample image corresponding to each original sample image into the teacher network model and the student network model respectively, and obtain the output of the teacher network model. k The first predicted image of each reconstructed sample image corresponding to each original sample image in the training and the first predicted image output by the student network model k The second predicted image of each reconstructed sample image corresponding to each original sample image in the training.

[0079] Specifically, for the Original sample images The corresponding Reconstructed sample image , in In the first training session, Original sample images The corresponding Reconstructed sample image After inputting the teacher network model and the student network model respectively, the output of the teacher network model can be obtained. k In the training Original sample images The corresponding Reconstructed sample image The first predicted image and the first output of the student network model k In the training Original sample images The corresponding Reconstructed sample image The second predicted image.

[0080] It should be noted that Original sample images The corresponding Reconstructed sample image The first predicted image can be the first Original sample images The corresponding Reconstructed sample image . No. Original sample images The corresponding Reconstructed sample image The second predicted image can also be the first Original sample images The corresponding Reconstructed sample image .

[0081] Using the reverse process of ROI Align technology, k The first predicted image of each reconstructed sample image corresponding to each original sample image in the training is fused to obtain the first k The fused prediction image corresponding to each original sample image in the training.

[0082] It should be noted that the inverse process of ROI Align technology is usually called Inverse ROI Align or ROIUnalign, and its goal is to remap the feature map processed by ROI Align technology back to the original image space.

[0083] The first In the training Original sample images After the first predicted image of each reconstructed sample image is input into the reverse ROI Align network, the reverse ROI Align network can k In the training Original sample images The corresponding reconstructed sample images are fused and the k In the training Original sample images The overlapping areas in the corresponding reconstructed sample images are fused by predictive average, and then the size of the output of the above reverse ROI Align network can be obtained as No. k In the training Original sample images The corresponding fused predicted image.

[0084] The mathematical expression of the calculation process of the reverse ROI Align network is as follows:

[0085] in, Indicated in In the first training session, Original sample images The corresponding Reconstructed sample image After inputting the teacher network model, the teacher network model outputs k In the training Original sample images The corresponding Reconstructed sample image The first predicted image of represents the mask function, which is obtained based on the sub-region bounding box; Indicates k In the training Original sample images The corresponding fused predicted image.

[0086] Based on k The fused prediction image corresponding to each original sample image in the training is obtained by using ROI Align technology. k Each secondary reconstructed image corresponding to each original sample image in the training is used as the k Each pseudo label corresponding to each original sample image in the training.

[0087] Specifically, obtain the k In the training Original sample images Corresponding fusion prediction image Afterwards, you can k In the training Original sample images Corresponding fusion prediction image Input the ROI Align network to obtain the output of the above ROI Align network. k In the training Original sample images The corresponding secondary reconstructed image.

[0088] It should be noted that, in the embodiment of the present invention, k In the training Original sample images The size of the corresponding secondary reconstructed image is k In the training Original sample images The corresponding Reconstructed sample image The same size as the first predicted image.

[0089] Get the output of the above ROI Align network k In the training Original sample images After each corresponding secondary reconstruction image, the output of the ROI Align network can be k In the training Original sample images Each corresponding secondary reconstructed image is determined as k In the training Original sample images Each corresponding pseudo label.

[0090] The embodiment of the present invention reconstructs multiple reconstructed sample images with different local spatial perspectives of the target area based on ROI Align technology, and then inputs the above-mentioned reconstructed sample images into the student network model and the teacher network model respectively. The output of the teacher network model is subjected to predicted average fusion and label realignment to obtain semantically consistent pseudo labels under different spatial perspectives, so as to perform spatial multi-perspective consistency training on the student network model.

[0091] As an optional embodiment, based on Each pseudo label corresponding to each original sample image in the training is constructed k The spatiotemporal consistency loss function corresponding to the training includes: k In the training, each original sample image is input into the final target history teacher network model to obtain the output of the final target history teacher network model. k The third predicted image corresponding to each original sample image in the training.

[0092] Among them, the last target history teacher network model is the target history teacher network model ranked last in the history teacher network model queue.

[0093] The history teacher network model queue is obtained based on the following steps: when it is determined that the first training is completed, the teacher network model in the first training is determined as the target history teacher network model corresponding to the first training, and the target history teacher network model corresponding to the first training is added to the history teacher network model queue.

[0094] In determining the The training is completed and confirmed When the difference between the number of training times corresponding to the target history teacher network model ranked first in the history teacher network model queue is the preset value, the first The teacher network model in the training is determined as the target historical teacher network model. .

[0095] When the number of target history teacher networks in the history teacher network model queue is less than the number threshold, The target history teacher network model corresponding to the training is inserted into the first place of the history teacher network model queue. When the number of target history teacher networks in the history teacher network model queue is equal to the quantity threshold, the first The target history teacher network model corresponding to the training is inserted into the first position of the history teacher network model queue, and the target history teacher network model at the last position in the history teacher network model queue is removed.

[0096] Specifically, the quantity threshold in the embodiment of the present invention is And the default value It can be determined based on prior knowledge and / or actual conditions. And the default value There is no restriction on the specific value of .

[0097] Based on The third predicted image and the first k The fused prediction image corresponding to each original sample image in the training is calculated to obtain the k The consistency weight map corresponding to each original sample image in the training.

[0098] Specifically, based on k The third predicted image and the first k The first predicted image of each reconstructed sample image corresponding to each original sample image in the training is calculated to obtain the first k In the training Original sample images The corresponding consistency weight graph is expressed as follows:

[0099] in, Indicates k In the training Original sample images The corresponding consistency weight graph; Indicates k In the training Original sample images a first predicted image corresponding to each reconstructed sample image; The output of the last target history teacher network model is Original sample images a corresponding third predicted image; express Function calculation; Indicates KL divergence (Kullback-Leibler Divergence) calculation.

[0100] Based on The consistency weight map corresponding to each original sample image in the training Each pseudo label corresponding to each original sample image in the training is constructed k The spatiotemporal consistency loss function corresponding to the training.

[0101] Specifically, based on k In the training Original sample images Corresponding consistency weight graph , can be k In the training Original sample images Each corresponding pseudo label is assigned a weight, and then the weighted k In the training Original sample images For each pseudo label, construct k The spatiotemporal consistency loss function corresponding to the training , the specific formula is as follows:

[0102] in, Indicates k In the training Original sample images Corresponding consistency weight graph The coordinates are The weight value of the point; Indicates In the training Original sample images Corresponding fusion prediction image The coordinates are The pixel value of the point; Indicated in In the first training session, Original sample images The corresponding Reconstructed sample image After inputting the student network model, the student network model outputs In the training Original sample images The corresponding Reconstructed sample image The second predicted image of Represents the Softmax function.

[0103] The embodiment of the present invention maintains a queue of historical teacher network models, performs consistency evaluation on the output predictions of the teacher network model currently in training and the last target historical teacher network model at the end of the queue, obtains a consistency weight map under different time-phase model states, re-weights the pseudo-labels, and thereby performs time-phase dynamic consistency training on the student network model by constructing a consistency loss.

[0104] As an optional embodiment, based on each reconstructed sample image, using the The teacher network model and student network model in the training are obtained. Each pseudo label corresponding to each original sample image in the training The embedded features corresponding to each reconstructed sample image and the embedded features corresponding to each original sample image in the training include: In the training, after inputting each reconstructed sample image corresponding to each original sample image into the student network model, the output of the feature extractor in the student network model is obtained. In the training, each original sample image corresponds to each reconstructed sample image, and in the In the training, after each original sample image is input into the final target history teacher network model, the feature extractor output of the final target history teacher network model is obtained. The feature image corresponding to each original sample image in the training.

[0105] It can be understood that the source domain model in the embodiment of the present invention includes a feature extractor and a classifier, and the student network model has the same network architecture as the source domain model, so the student network model and the teacher network model in the embodiment of the present invention both include a feature extractor and a classifier.

[0106] In the In the first training session, Original sample images The corresponding Reconstructed sample image After inputting the student network model, we can get the output of the feature extractor in the student network model. In the training Original sample images The corresponding Reconstructed sample image The corresponding feature image.

[0107] In the In the first training session, Original sample images After inputting the final target history teacher network model, the feature extractor output of the final target history teacher network model can be obtained. k In the training Original sample images The corresponding feature image.

[0108] The first k In the training, the feature image corresponding to each reconstructed sample image and the feature image corresponding to each original sample image are projected into the feature space of preset dimensions respectively to obtain the first k The embedded features corresponding to each reconstructed sample image and the embedded features corresponding to each original sample image in the training are obtained.

[0109] It should be noted that the preset dimensions in the embodiments of the present invention may be determined based on prior knowledge and / or actual conditions, and the specific values ​​of the preset dimensions in the embodiments of the present invention are not limited.

[0110] Optionally, the preset dimension in the embodiment of the present invention may be 128.

[0111] Get the k Each original sample image in the training The corresponding feature image of each reconstructed sample image and the k After each feature image corresponding to the original sample image in the training, a projection head consisting of two layers of multilayer perceptron (MLP) can be used to respectively k The feature image corresponding to each reconstructed sample image corresponding to each original sample image in the training and the k The feature image corresponding to each original sample image in the training is projected into a 128-dimensional feature space, thereby obtaining the first k The embedding features corresponding to each reconstructed sample image corresponding to each original sample image in the training The embedded features corresponding to each original sample image .

[0112] It should be noted that the encoded embedding is the feature map output by the model's encoder (feature extractor). The output of the model contains semantic categories. After resampling the feature map and the model's output to the same size, feature areas of the same semantic category and feature areas of different semantic categories can be obtained. In this way, positive samples, negative samples, and anchor points can be divided.

[0113] Get the k The embedding features corresponding to each reconstructed sample image corresponding to each original sample image in the training The embedded features corresponding to each original sample image Afterwards, based on k The embedding features corresponding to each reconstructed sample image corresponding to each original sample image in the training The embedded features corresponding to each original sample image Construct anchor points, positive samples, and negative samples.

[0114] The anchor point is the encoding embedding of the student network model, the positive sample is the encoding embedding with the same semantic category as the anchor point, and the encoding embedding of other different semantic categories is the negative sample. The spatiotemporal contrast loss function corresponding to the training is expressed as follows:

[0115] in, Indicates The value of the spatiotemporal contrast loss function corresponding to the training; Indicates The embedding features corresponding to each reconstructed sample image corresponding to each original sample image in the training The semantic category of Indicates The embedded features corresponding to each original sample image in the training The semantic category of Represents the number of embedded elements, where an embedded element refers to the pixel unit of the feature map; represents the temperature coefficient; represents a positive integer; Indicates The embedding vector corresponding to the embedding elements.

[0116] The embodiment of the present invention obtains embedded features by utilizing the feature extractor of the student network model, and obtains feature images by utilizing the final target history teacher network model. The embedded features of the student network model are used as anchor samples, the encoded embeddings projected by the pseudo-labels with the same semantic category as the anchor samples are used as positive samples, and the encoded embeddings belonging to other semantic categories are used as negative samples. The spatiotemporal contrast loss is constructed to perform spatiotemporal contrast training on the student network model.

[0117] Based on The spatiotemporal consistency loss function and the spatiotemporal contrast loss function corresponding to the training are combined with the minimum entropy loss function to construct the The target loss function corresponding to the training.

[0118] Specifically, no. The minimum entropy loss function corresponding to the training The formula is as follows:

[0119] in, Indicates semantic categories; Represents the sum of semantic categories.

[0120] No. The formula of the objective loss function corresponding to the training is as follows:

[0121] in, Indicates The target loss function value corresponding to the training; and Set to the relative contribution of the spatiotemporal contrast loss function and the minimum entropy loss function in the model gradient update optimization process.

[0122] Calculate the The function value of the target loss function corresponding to the training is based on the The function value of the target loss function corresponding to the training time determines whether the student network model has not converged or When the number of training times is not greater than the maximum, the model parameters of the student network model are updated. Increase by 1 and return to execute in k In the training, based on each reconstructed sample image, using the k The teacher network model and student network model in the training are obtained. k Each pseudo label corresponding to each original sample image in the training k In the steps of embedding features corresponding to each reconstructed sample image and embedding features corresponding to each original sample image in the training, based on the The function value of the target loss function corresponding to the training determines whether the student network model converges or When the number of training times is greater than the maximum number, it is determined that the student network model has been trained, and the trained student network model is determined as the remote sensing extraction model of the land features corresponding to the target area.

[0123] Step 104: Based on the remote sensing extraction model of the ground object elements corresponding to the target area and the remote sensing image of the target area, the remote sensing extraction result of the ground object elements of the target area is obtained.

[0124] As an optional embodiment, based on the remote sensing extraction model of the land features corresponding to the target area and the remote sensing image of the target area, the remote sensing extraction results of the land features in the target area are obtained, including: inputting each original sample image into the remote sensing extraction model of the land features corresponding to the target area, and obtaining the predicted image corresponding to each original sample image output by the remote sensing extraction model of the land features corresponding to the target area.

[0125] The predicted images corresponding to the original sample images are spliced ​​to obtain the predicted image of the target area.

[0126] Information is extracted from the predicted image of the target area to obtain the remote sensing extraction results of the ground objects in the target area.

[0127] The embodiment of the present invention obtains a remote sensing image of the target area as the target domain image, and then obtains multiple reconstructed sample images with different local spatial perspectives of the target area based on the target domain image. Based on the reconstructed sample images and the teacher network model, the student network model is improved by domain transfer training to obtain the trained student network model as the remote sensing extraction model of the land object elements corresponding to the target area. Then, based on the remote sensing extraction model of the land object elements corresponding to the target area and the remote sensing image of the target area, the remote sensing extraction result of the land object elements of the target area is obtained. In the case where the training sample data used to train the source domain model is missing, the source domain model can be domain transferred through the improved domain transfer training, and the features in the remote sensing image of the target area can be transferred to the source domain model, which can significantly improve the adaptability of the source domain model to the target area, thereby improving the accuracy of remote sensing extraction of land object elements in the target area.

[0128] The flowchart of the domain migration method for remote sensing extraction of regional surface elements provided by the present invention is different from the traditional domain migration technology of the related art, which only considers the output information of the current model iteration to define the reliability of the pseudo-label. The improved domain migration method in the present invention includes two parts: spatial multi-view consistency learning and temporal dynamic consistency learning. Spatial multi-view consistency learning improves the contextual consistency and reliability of pseudo-labels through the fusion alignment of multi-view enhancement and output prediction, and temporal dynamic consistency learning constrains the consistency of current and historical model knowledge to regularize the evolution direction of the model, preventing the model from gradually forgetting knowledge due to lack of supervision information.

[0129] In addition, the domain migration method for regional surface element remote sensing extraction provided by the present invention designs a new spatiotemporal comparison strategy, which is different from the traditional domain migration technology in the related technology that lacks semantic association constraints at the spatiotemporal level. The spatiotemporal comparison strategy in the present invention embeds the intrinsic semantic associations of different category features through models at different times and different spatial perspectives, so as to improve the category discrimination representation of the surface element remote sensing extraction classification model in the target domain, thereby improving the generalization ability of the model.

[0130] Compared with the related art, the domain migration method for regional surface element remote sensing extraction provided by the present invention solves the prediction inconsistency problem of the surface element remote sensing extraction classification model in different spatial contexts through a spatial multi-perspective consistency learning mechanism. Through the spatial multi-perspective enhancement and prediction fusion process, a more stable and reliable multi-perspective unified pseudo-label can be obtained, so that the stability and adaptability of the domain migration model are stronger; the knowledge forgetting problem caused by the lack of supervisory information guidance in the domain migration process is solved through a temporal dynamic consistency learning mechanism, and the model knowledge of historical moments is introduced to regularize the training direction of the model, so that the model has the ability to self-correct the current pseudo-label noise, so that the model's element extraction and classification results in the target domain are more complete and the false detection rate is lower; through a spatiotemporal contrast learning strategy, the cognitive bias problem caused by domain shift in large-area surface element extraction and classification tasks is solved, and contrast learning is used to constrain the intrinsic semantic association of different category features, so that the model has a stronger ability to distinguish different surface elements with small feature differences.

[0131] Table 1 is one of the accuracy comparison tables between the domain migration method for remote sensing extraction of regional surface elements provided by the present invention and the traditional domain migration method in the related art. Figure 5 This is one of the effect comparison diagrams of the domain migration method for regional surface element remote sensing extraction provided by the present invention and the traditional domain migration method for regional surface element remote sensing extraction in the related art.

[0132] The domain migration method for regional surface element remote sensing extraction provided by the present invention is applied to the loveDA urban-rural domain adaptation dataset, and the application effect is compared with traditional domain migration methods such as HCL, IAPC and DT-ST. The obtained accuracy evaluation results and effect evaluation results are shown in Tables 1 and 2. Figure 5 shown.

[0133] Table 1 Comparison table 1 of the accuracy of the domain migration method for regional surface element remote sensing extraction provided by the present invention and the traditional domain migration method in the related art

[0134] As shown in Table 1 and Figure 5 As shown, compared with the traditional domain migration method in the related art, the domain migration method for regional surface element remote sensing extraction provided by the present invention can significantly improve the accuracy of surface element remote sensing extraction.

[0135] Table 2 is the second comparison table of the accuracy between the domain migration method for remote sensing extraction of regional surface elements provided by the present invention and the traditional domain migration method in the related art. Figure 6This is the second comparison diagram of the effects of the domain migration method for remote sensing extraction of regional surface elements provided by the present invention and the traditional domain migration method for remote sensing extraction of regional surface elements in the related art.

[0136] The domain migration method for regional surface feature remote sensing extraction provided by the present invention is further applied to a large-scale surface feature extraction classification dataset in a large area, and the application effect is compared with traditional domain migration methods such as HCL, IAPC and DT-ST. The obtained accuracy evaluation results and effect evaluation results are shown in Table 2 and Table 3. Figure 6 shown.

[0137] Table 2 Comparison table 2 of the accuracy of the domain migration method for regional surface element remote sensing extraction provided by the present invention and the traditional domain migration method in the related art

[0138] As shown in Table 2 and Figure 6 As shown, compared with the traditional domain migration method in the related art, the domain migration method for regional surface element remote sensing extraction provided by the present invention can significantly improve the accuracy of surface element remote sensing extraction.

[0139] Figure 7 This is a schematic diagram of the structure of the domain migration device for remote sensing extraction of regional surface elements provided by the present invention. Figure 7 The domain migration device for remote sensing extraction of regional surface elements provided by the present invention is described. The domain migration device for remote sensing extraction of regional surface elements described below and the domain migration method for remote sensing extraction of regional surface elements provided by the present invention described above can be referred to each other. Figure 7 As shown, the device includes: a data acquisition module 701, a sample reconstruction module 702, a model training module 703 and a remote sensing extraction module 704.

[0140] The data acquisition module 701 is used to acquire a remote sensing image of a target area as a target domain image.

[0141] The sample reconstruction module 702 is used to obtain a plurality of reconstructed sample images with different local spatial perspectives of the target area based on the target domain image.

[0142] The model training module 703 is used to perform improved domain transfer training on the student network model based on each reconstructed sample image and the teacher network model to obtain a trained student network model as the remote sensing extraction model of the ground object elements corresponding to the target area. The model structure of the teacher network model and the student network model is the same as the model structure of the source domain model. The initial model parameters of the teacher network model and the student network model are the model parameters of the source domain model. The source domain model is trained based on the remote sensing image of the sample area and the remote sensing image of the labeled sample area. The annotation information in the remote sensing image of the labeled sample area is used to indicate at least one of the type, identification, quantity and area range of the ground object elements.

[0143] The remote sensing extraction module 704 is used to obtain the remote sensing extraction results of the ground object elements in the target area based on the ground object element remote sensing extraction model corresponding to the target area and the remote sensing image of the target area.

[0144] Specifically, the data acquisition module 701, the sample reconstruction module 702, the model training module 703 and the remote sensing extraction module 704 are electrically connected.

[0145] The domain migration device for remote sensing extraction of regional surface elements in the embodiment of the present invention obtains a remote sensing image of the target area as the target domain image, and then obtains multiple reconstructed sample images with different local spatial perspectives of the target area based on the target domain image. Based on each reconstructed sample image and the teacher network model, the student network model is improved by domain migration training to obtain the trained student network model as the remote sensing extraction model of the object elements corresponding to the target area, and then based on the remote sensing extraction model of the object elements corresponding to the target area and the remote sensing image of the target area, the remote sensing extraction result of the object elements of the target area is obtained. In the case where the training sample data used to train the source domain model is missing, the source domain model can be domain migrated through the improved domain migration training, and the features in the remote sensing image of the target area can be migrated to the source domain model, which can significantly improve the adaptability of the source domain model to the target area, thereby improving the accuracy of remote sensing extraction of object elements in the target area.

[0146] Figure 8 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 8As shown, the electronic device may include: a processor (processor) 810, a communication interface (Communications Interface) 820, a memory (memory) 830 and a communication bus 840, wherein the processor 810, the communication interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 can call the logic instructions in the memory 830 to execute a domain migration method for remote sensing extraction of regional surface elements, the method comprising: obtaining a remote sensing image of the target area as a target domain image; based on the target domain image, obtaining multiple reconstructed sample images with different local spatial perspectives of the target area; based on each reconstructed sample image and the teacher network model, performing improved domain migration training on the student network model to obtain a trained student network model as a remote sensing extraction model of the surface elements corresponding to the target area, the model structure of the teacher network model and the student network model is the same as the model structure of the source domain model, the initial model parameters of the teacher network model and the student network model are the model parameters of the source domain model, the source domain model is obtained by training based on the remote sensing image of the sample area and the remote sensing image of the sample area after annotation, and the annotation information in the remote sensing image of the sample area after annotation is used to indicate at least one of the type, identification, quantity and regional range of the surface elements; based on the remote sensing extraction model of the surface elements corresponding to the target area and the remote sensing image of the target area, the remote sensing extraction result of the surface elements of the target area is obtained.

[0147] In addition, the logic instructions in the above-mentioned memory 830 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0148] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the domain migration method for remote sensing extraction of regional surface elements provided by the above methods, the method including: obtaining a remote sensing image of the target area as a target domain image; based on the target domain image, obtaining multiple reconstructed sample images with different local spatial perspectives of the target area; based on each reconstructed sample image and the teacher network model, performing improved domain migration training on the student network model to obtain a trained student network model as a remote sensing extraction model of land features corresponding to the target area, the model structure of the teacher network model and the student network model is the same as the model structure of the source domain model, the initial model parameters of the teacher network model and the student network model are the model parameters of the source domain model, the source domain model is obtained by training based on the remote sensing image of the sample area and the remote sensing image of the sample area after annotation, and the annotation information in the remote sensing image of the sample area after annotation is used to indicate at least one of the type, identification, quantity and regional range of the land features; based on the remote sensing extraction model of the land features corresponding to the target area and the remote sensing image of the target area, the remote sensing extraction result of the land features of the target area is obtained.

[0149] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the domain migration method for remote sensing extraction of regional surface elements provided by the above-mentioned methods, the method comprising: obtaining a remote sensing image of a target area as a target domain image; based on the target domain image, obtaining a plurality of reconstructed sample images with different local spatial perspectives of the target area; based on each reconstructed sample image and a teacher network model, performing improved domain migration training on a student network model to obtain a trained student network model as a remote sensing extraction model of land features corresponding to the target area, the model structure of the teacher network model and the student network model being the same as the model structure of a source domain model, the initial model parameters of the teacher network model and the student network model being the model parameters of a source domain model, the source domain model being obtained by training based on a remote sensing image of a sample area and a remote sensing image of an annotated sample area, the annotation information in the remote sensing image of an annotated sample area being used to indicate at least one of the type, identification, quantity and regional range of land features; based on the remote sensing extraction model of land features corresponding to the target area and the remote sensing image of the target area, obtaining a remote sensing extraction result of land features of the target area.

[0150] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0151] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0152] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A domain migration method for remote sensing extraction of regional surface elements, characterized in that: include: Acquire a remote sensing image of the target area as the target domain image; Based on the target domain image, a plurality of reconstructed sample images having different local spatial perspectives of the target area are obtained; Based on the reconstructed sample images and the teacher network model, an improved domain transfer training is performed on the student network model to obtain a trained student network model as a remote sensing extraction model of land features corresponding to the target area, the model structures of the teacher network model and the student network model are the same as the model structures of the source domain model, the initial model parameters of the teacher network model and the student network model are the model parameters of the source domain model, the source domain model is obtained by training based on the remote sensing image of the sample area and the remote sensing image of the labeled sample area, and the annotation information in the remote sensing image of the labeled sample area is used to indicate at least one of the type, identification, quantity and area range of the land features; Based on the remote sensing extraction model of the ground object elements corresponding to the target area and the remote sensing image of the target area, the remote sensing extraction result of the ground object elements of the target area is obtained.

2. The domain migration method for remote sensing extraction of regional surface elements according to claim 1 is characterized in that: The step of obtaining a plurality of reconstructed sample images having different local spatial perspectives of the target area based on the target domain image includes: Based on a sliding window of a preset size and a predefined overlap rate, the target domain image is cropped to obtain a plurality of original sample images; For each original sample image, randomly generate a plurality of sub-region boundary boxes of the same size, located in different regions of each original sample image and overlapping with each other in the original sample image; Based on the bounding boxes of each sub-region in each original sample image and the bounding box of each original sample image, a region of interest alignment (ROI Align) technology is used to obtain each reconstructed sample image corresponding to each original sample image.

3. The domain migration method for remote sensing extraction of regional surface elements according to claim 2 is characterized in that: The improved domain transfer training is performed on the student network model based on the reconstructed sample images and the teacher network model to obtain a trained student network model as a remote sensing extraction model of ground features corresponding to the target area, including: In the k In the training, based on the reconstructed sample images, the k The teacher network model and the student network model in the training are obtained. k Each pseudo label corresponding to each original sample image in the training k The embedded features corresponding to each reconstructed sample image corresponding to each original sample image in the training and the embedded features corresponding to each original sample image; Based on k Each pseudo label corresponding to each original sample image in the training is constructed k The spatiotemporal consistency loss function corresponding to the training is based on the The embedded features corresponding to each reconstructed sample image corresponding to each original sample image in the training and the embedded features corresponding to each original sample image are constructed to obtain the first k The spatiotemporal contrast loss function corresponding to the training; Based on The spatiotemporal consistency loss function and the spatiotemporal contrast loss function corresponding to the training are combined with the minimum entropy loss function to construct the k The target loss function corresponding to the training; Calculate the The function value of the target loss function corresponding to the training is based on the k The function value of the target loss function corresponding to the training time determines whether the student network model has not converged or k When the number of training times is not greater than the maximum, the model parameters of the student network model are updated. k Increase by 1 and return to execute the k In the training, based on the reconstructed sample images, the k The teacher network model and the student network model in the training are obtained. k Each pseudo label corresponding to each original sample image in the training k The steps of embedding features corresponding to each reconstructed sample image corresponding to each original sample image and embedding features corresponding to each original sample image in the first training are based on the first k The function value of the target loss function corresponding to the training time determines whether the student network model converges or k When the number of training times is greater than the maximum number, it is determined that the student network model has been trained, and the trained student network model is determined as the remote sensing extraction model of the land feature corresponding to the target area.

4. The domain migration method for remote sensing extraction of regional surface elements according to claim 3 is characterized in that: Based on each of the reconstructed sample images, using the k The teacher network model and the student network model in the training are obtained. k Each pseudo label corresponding to each original sample image in the training includes: In the k In the training, the model parameters of the teacher network model are updated based on the exponential moving average method; The reconstructed sample images corresponding to each original sample image are input into the teacher network model and the student network model respectively, and the first k The first predicted image of each reconstructed sample image corresponding to each original sample image in the training and the first predicted image output by the student network model k A second predicted image of each reconstructed sample image corresponding to each original sample image in the training; Using the reverse process of ROI Align technology, k The first predicted image of each reconstructed sample image corresponding to each original sample image in the training is fused to obtain the first k The fused prediction image corresponding to each original sample image in the training; Based on k The fused prediction image corresponding to each original sample image in the training is obtained by using ROI Align technology. k Each secondary reconstructed image corresponding to each original sample image in the training is used as the k Each pseudo label corresponding to each original sample image in the training.

5. The domain migration method for remote sensing extraction of regional surface elements according to claim 3, characterized in that: The k Each pseudo label corresponding to each original sample image in the training is constructed k The spatiotemporal consistency loss function corresponding to the training includes: In the k In the training, each original sample image is input into the final target history teacher network model to obtain the output of the final target history teacher network model. k The third predicted image corresponding to each original sample image in the training; Based on k The third predicted image and the first k The fused prediction image corresponding to each original sample image in the training is calculated to obtain the k A consistency weight map corresponding to each original sample image in the training; Based on k The consistency weight map corresponding to each original sample image in the training and k Each pseudo label corresponding to each original sample image in the training is constructed to obtain the k The spatiotemporal consistency loss function corresponding to the training; The last target history teacher network model is the last target history teacher network model in the history teacher network model queue; The history teacher network model queue is obtained based on the following steps: When it is determined that the first training is finished, the teacher network model in the first training is determined as the target history teacher network model corresponding to the first training, and the target history teacher network model corresponding to the first training is added to the history teacher network model queue; In determining the m The training is completed and confirmed m When the difference between the number of training times corresponding to the target history teacher network model ranked first in the history teacher network model queue is a preset value, the first m The teacher network model in the training is determined as the target historical teacher network model. ; When the number of target history teacher networks in the history teacher network model queue is less than the number threshold, m The target history teacher network model corresponding to the training is inserted into the first position of the history teacher network model queue. When the number of target history teacher networks in the history teacher network model queue is equal to the number threshold, the first m The target history teacher network model corresponding to the training is inserted into the first position of the history teacher network model queue, and the target history teacher network model at the last position in the history teacher network model queue is removed.

6. The domain migration method for remote sensing extraction of regional surface elements according to claim 3, characterized in that: Based on each of the reconstructed sample images, using the k The teacher network model and the student network model in the training are obtained. k Each pseudo label corresponding to each original sample image in the training k The embedded features corresponding to each reconstructed sample image corresponding to each original sample image in the training and the embedded features corresponding to each original sample image include: In the k In the training, after inputting each reconstructed sample image corresponding to each original sample image into the student network model, the output of the feature extractor in the student network model is obtained. In the training, each original sample image corresponds to each reconstructed sample image corresponding to the feature image, k In the training, after each original sample image is input into the final target history teacher network model, the feature extractor output of the final target history teacher network model is obtained. k The feature image corresponding to each original sample image in the training; The first k The feature image corresponding to each reconstructed sample image corresponding to each original sample image and the feature image corresponding to each original sample image in the training are projected into a feature space of a preset dimension respectively to obtain the first k The embedded features corresponding to each reconstructed sample image corresponding to each original sample image in the training and the embedded features corresponding to each original sample image.

7. The domain migration method for remote sensing extraction of regional surface elements according to any one of claims 1 to 6, characterized in that: The step of obtaining a remote sensing extraction result of a ground object element of the target area based on the ground object element remote sensing extraction model corresponding to the target area and the remote sensing image of the target area includes: Input each original sample image into the remote sensing extraction model of the ground feature elements corresponding to the target area, and obtain the predicted image corresponding to each original sample image output by the remote sensing extraction model of the ground feature elements corresponding to the target area; splicing the predicted images corresponding to the original sample images to obtain the predicted image of the target area; Information is extracted from the predicted image of the target area to obtain remote sensing extraction results of ground objects in the target area.

8. A domain migration device for remote sensing extraction of regional surface elements, characterized in that: include: A data acquisition module is used to acquire remote sensing images of the target area as target domain images; A sample reconstruction module, used to obtain a plurality of reconstructed sample images with different local spatial perspectives of the target area based on the target domain image; A model training module, for performing improved domain transfer training on the student network model based on each of the reconstructed sample images and the teacher network model, to obtain a trained student network model as a remote sensing extraction model of land features corresponding to the target area, wherein the model structures of the teacher network model and the student network model are the same as the model structure of the source domain model, and the initial model parameters of the teacher network model and the student network model are the model parameters of the source domain model, and the source domain model is obtained by training based on the remote sensing image of the sample area and the remote sensing image of the labeled sample area, and the annotation information in the remote sensing image of the labeled sample area is used to indicate at least one of the type, identification, quantity and area range of the land features; The remote sensing extraction module is used to obtain the remote sensing extraction results of the ground object elements in the target area based on the remote sensing extraction model of the ground object elements corresponding to the target area and the remote sensing image of the target area.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the domain migration method for remote sensing extraction of regional surface elements as described in any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the domain migration method for remote sensing extraction of regional surface elements as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Remote sensing image domain adaptive retrieval method based on pseudo-label consistency learning

    CN115292532A

  • Ground feature element segmentation model training method and ground feature element segmentation method and device

    CN116977633A

  • Remote sensing image cross-domain small sample classification method based on pseudo label uncertainty perception

    CN117152503A

  • Object detection model training method, detection method, apparatus, device and medium

    WO2024120157A1

Cited By

  • Complex farming area idle cultivated land extraction method and device, model training method and device, storage medium and terminal

    CN120236096A

  • Method and system for automatically extracting topographic elements based on deep learning

    CN120853035A