Domain Migration Method, Device, Equipment and Medium for Remote Sensing Extraction of Regional Surface Elements
By generating multi-view reconstruction sample images in the target area and combining improved training of the teacher-student network model, the model is optimized to adapt to the difference in feature distribution, and the problem of inaccurate remote sensing extraction of land objects in traditional domain migration methods is solved, and a higher extraction accuracy is achieved.
Patent Information
- Application Number
- CN202510437532.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-04-09
AI Technical Summary
In the prior art, when there are differences in the characteristic distribution of training sample data and test data, it is difficult to accurately perform remote sensing extraction of land objects. The model optimization effect is poor, resulting in low accuracy of remote sensing extraction of land objects objects.
By obtaining remote sensing images of the target area, multiple reconstructed sample images from different local spatial perspectives are generated, and improved domain migration training is used to combine the space-time consistency loss function and dynamic teacher-student mechanism to optimize the student network model and improve the model's adaptability in the target area.
The accuracy of remote sensing extraction of land elements in the target area of the source domain model has been significantly improved, the model's adaptability under the different characteristics of the characteristics, and the accuracy of remote sensing extraction of land elements has been improved.
Smart Images

Figure CN119964163B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a domain migration method, device, equipment and medium for remote sensing extraction of regional surface elements. Background Art
[0002] Remote sensing extraction of land features refers to the process of using remote sensing images to obtain information such as the number and spatial distribution of land features such as buildings, roads, water bodies, cultivated land and forest land on the earth's surface. Remote sensing extraction of land features has been widely used in many fields such as resource management, environmental testing, urban planning and disaster monitoring.
[0003] The traditional domain transfer method in the related art can use the model trained based on the training sample data and sample labels to realize the remote sensing extraction of ground features of the test data. However, due to the influence of spatiotemporal heterogeneity, when the training sample data and the test data are respectively obtained based on remote sensing data corresponding to different regions, different times or different surface morphologies, there are differences in feature distribution between the training sample data and the test data, which do not obey the independent and identically distributed assumption of machine learning, resulting in low accuracy of remote sensing extraction of ground features of the test data based on the above model, and problems such as insufficient model generalization.
[0004] Domain migration technology can migrate the knowledge of training sample data to test data in scenarios where there are differences in feature distribution between training sample data and test data (domain shift) and training sample data is missing, relying only on models trained with training sample data and unlabeled test data, so as to improve the generalization ability of the model in test data without adding additional annotation costs. However, the optimization effect of optimizing the trained model based on traditional domain migration technology in related technologies is not good. Based on the optimized model, it is still difficult to accurately perform remote sensing extraction of ground object elements on test data when there are differences in feature distribution between training sample data and test data. Therefore, how to better optimize the trained model, so as to more accurately perform remote sensing extraction of ground object elements on test data when there are differences in feature distribution between training sample data and test data, is a technical problem that needs to be solved urgently in this field. Summary of the invention
[0005] The present invention provides a domain migration method, device, equipment and medium for remote sensing extraction of regional surface elements, which is used to solve the defects that the traditional domain migration technology in the prior art has a poor optimization effect on the trained model, and it is still difficult to accurately perform remote sensing extraction of land object elements on the test data based on the optimized model when there is a difference in feature distribution between the training sample data and the test data, so as to achieve better optimization of the trained model, thereby more accurately performing remote sensing extraction of land object elements on the test data when there is a difference in feature distribution between the training sample data and the test data.
[0006] The present invention provides a domain transfer method for remote sensing extraction of regional surface elements, including the following steps.
[0007] Obtain a remote sensing image of the target area as the target domain image; based on the target domain image, obtain multiple reconstructed sample images with different local spatial perspectives of the target area; based on each of the reconstructed sample images and a teacher network model, perform improved domain transfer training on a student network model to obtain a trained student network model as the remote sensing extraction model of the ground feature elements corresponding to the target area. The model structures of the teacher network model and the student network model are the same as the model structure of the source domain model, and the initial model parameters of the teacher network model and the student network model are the model parameters of the source domain model. The source domain model is trained based on the remote sensing image of the sample area and the remote sensing image of the labeled sample area, and the labeling information in the remote sensing image of the labeled sample area is used to indicate at least one of the type, identifier, quantity, and regional scope of the ground feature elements; based on the remote sensing extraction model of the ground feature elements corresponding to the target area and the remote sensing image of the target area, obtain the remote sensing extraction result of the ground feature elements of the target area.
[0008] According to the domain transfer method for remote sensing extraction of regional surface elements provided by the present invention, the step of obtaining multiple reconstructed sample images with different local spatial perspectives of the target area based on the target domain image includes: cropping the target domain image based on a sliding window of a preset size and a predefined overlap rate to obtain multiple original sample images; for each original sample image, randomly generate multiple sub-region bounding boxes with the same size, located in different regions of each original sample image and overlapping with each other; based on each sub-region bounding box and the bounding box of each original sample image in each original sample image, use the Region of Interest Align (ROIAlign) technique to obtain each reconstructed sample image corresponding to each original sample image.
[0009] According to the domain transfer method for remote sensing extraction of regional surface elements provided by the present invention, the step of performing improved domain transfer training on a student network model based on each of the reconstructed sample images and a teacher network model to obtain a trained student network model as the remote sensing extraction model of the ground feature elements corresponding to the target area includes:
[0010] In the k th training, based on each of the reconstructed sample images, use the teacher network model and the student network model in the k th training to obtain each pseudo-label corresponding to each original sample image in the k th training and the kThe embedding features corresponding to each reconstructed sample image corresponding to each original sample image in the current training and the embedding features corresponding to each original sample image;
[0011] Based on the pseudo-labels corresponding to each original sample image in the k th training, construct the spatio-temporal consistency loss function corresponding to the k th training, and based on the embedding features corresponding to each reconstructed sample image corresponding to each original sample image and the embedding features corresponding to each original sample image in the th training, construct the spatio-temporal contrast loss function corresponding to the k th training;
[0012] Based on the spatio-temporal consistency loss function and spatio-temporal contrast loss function corresponding to the th training, and in combination with the minimum entropy loss function, construct the target loss function corresponding to the k th training;
[0013] Calculate the function value of the target loss function corresponding to the th training. When it is determined that the student network model has not converged or k is not greater than the maximum number of training times based on the function value of the target loss function corresponding to the k th training, update the model parameters of the student network model, k increase by 1, and return to execute the step of obtaining the pseudo-labels corresponding to each original sample image in the k th training and the embedding features corresponding to each reconstructed sample image corresponding to each original sample image and the embedding features corresponding to each original sample image in the k th training based on each of the reconstructed sample images, using the teacher network model and the student network model in the k th training. When it is determined that the student network model has converged or k is greater than the maximum number of training times based on the function value of the target loss function corresponding to the k th training, determine that the student network model has been trained well, and determine the trained student network model as the remote sensing extraction model of the ground feature elements corresponding to the target area. k Based on a domain transfer method for remote sensing extraction of regional surface elements provided by the present invention, the step of obtaining the pseudo-labels corresponding to each original sample image in the
[0014] th training based on each of the reconstructed sample images, using the teacher network model and the student network model in the k th training includes: k
[0015] In the k th training, based on the exponential moving average method, update the model parameters of the teacher network model;
[0016] Input the reconstructed sample images corresponding to each original sample image into the teacher network model and the student network model respectively, and obtain the first predicted images of each reconstructed sample image corresponding to each original sample image in the k th training output by the teacher network model and the second predicted images of each reconstructed sample image corresponding to each original sample image in the k th training output by the student network model;
[0017] Using the reverse process of the ROI Align technology, perform data fusion on the first predicted images of the reconstructed sample images corresponding to each original sample image in the k th training, and obtain the fused predicted images of each original sample image in the k th training;
[0018] Based on the fused predicted images of each original sample image in the k th training, use the ROI Align technology to obtain each secondary reconstructed image corresponding to each original sample image in the k th training, and use it as each pseudo-label corresponding to each original sample image in the k th training.
[0019] According to a domain transfer method for remote sensing extraction of regional surface elements provided by the present invention, based on each pseudo-label corresponding to each original sample image in the k th training, construct the spatio-temporal consistency loss function corresponding to the k th training, including:
[0020] In the k th training, input each original sample image into the last target historical teacher network model, and obtain the third predicted images of each original sample image in the k th training output by the last target historical teacher network model;
[0021] Based on the third predicted images of each original sample image in the k th training and the fused predicted images of each original sample image in the k th training, calculate the consistency weight map of each original sample image in the k th training;
[0022] Based on the kThe consistency weight map corresponding to each original sample image in the training and k Each pseudo label corresponding to each original sample image in the training is constructed to obtain the k The spatiotemporal consistency loss function corresponding to the training;
[0023] The last target history teacher network model is the last target history teacher network model in the history teacher network model queue;
[0024] The history teacher network model queue is obtained based on the following steps:
[0025] When it is determined that the first training is finished, the teacher network model in the first training is determined as the target history teacher network model corresponding to the first training, and the target history teacher network model corresponding to the first training is added to the history teacher network model queue;
[0026] In determining the m The training is completed and confirmed m When the difference between the number of training times corresponding to the target history teacher network model ranked first in the history teacher network model queue is a preset value, the first m The teacher network model in the training is determined as the target historical teacher network model. ;
[0027] When the number of target history teacher networks in the history teacher network model queue is less than the number threshold, m The target history teacher network model corresponding to the training is inserted into the first position of the history teacher network model queue. When the number of target history teacher networks in the history teacher network model queue is equal to the number threshold, the first m The target history teacher network model corresponding to the training is inserted into the first position of the history teacher network model queue, and the target history teacher network model at the last position in the history teacher network model queue is removed.
[0028] According to a domain migration method for remote sensing extraction of regional surface elements provided by the present invention, based on each of the reconstructed sample images, using the k The teacher network model and the student network model in the training are obtained. k Each pseudo label corresponding to each original sample image in the training k The embedded features corresponding to each reconstructed sample image corresponding to each original sample image in the training and the embedded features corresponding to each original sample image include:
[0029] In the kIn the th training, after inputting each reconstructed sample image corresponding to each original sample image into the student network model, the feature image corresponding to each reconstructed sample image corresponding to each original sample image in the k th training output by the feature extractor in the student network model is obtained. In the k th training, after inputting each original sample image into the final target historical teacher network model, the feature image corresponding to each original sample image in the
[0030] th training output by the feature extractor in the final target historical teacher network model is obtained; k The feature images corresponding to each reconstructed sample image corresponding to each original sample image and the feature images corresponding to each original sample image in the k th training are respectively projected into a feature space with a preset dimension to obtain the embedded features corresponding to each reconstructed sample image corresponding to each original sample image and the embedded features corresponding to each original sample image in the
[0031] The present invention also provides a domain transfer device for remote sensing extraction of regional surface elements, including the following modules.
[0032] A data acquisition module, configured to acquire a remote sensing image of a target area as a target domain image; a sample reconstruction module, configured to perform improved domain transfer training on a student network model based on each of the reconstructed sample images and a teacher network model to obtain a trained student network model as a remote sensing extraction model of ground object elements corresponding to the target area. The model structures of the teacher network model and the student network model are the same as the model structure of the source domain model, and the initial model parameters of the teacher network model and the student network model are the model parameters of the source domain model. The source domain model is trained based on the remote sensing image of the sample area and the remote sensing image of the labeled sample area, and the labeling information in the remote sensing image of the labeled sample area is used to indicate at least one of the type, identifier, quantity, and regional range of the ground object elements; a remote sensing extraction module, configured to obtain a remote sensing extraction result of the ground object elements in the target area based on the remote sensing extraction model of the ground object elements corresponding to the target area and the remote sensing image of the target area.
[0033] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the domain transfer method for remote sensing extraction of regional surface elements as described in any one of the above is implemented.
[0034] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the domain transfer method for remote sensing extraction of regional surface elements as described in any one of the above is implemented.
[0035] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the domain transfer method for remote sensing extraction of regional surface elements as described in any one of the above is implemented.
[0036] The domain transfer method, device, equipment and medium for remote sensing extraction of regional surface elements provided by the present invention, after obtaining a remote sensing image of a target area as a target domain image, based on the target domain image, obtains multiple reconstructed sample images with different local spatial perspectives of the target area. Based on each reconstructed sample image and a teacher network model, an improved domain transfer training is performed on a student network model to obtain a trained student network model, which is used as a remote sensing extraction model for ground object elements corresponding to the target area. Then, based on the remote sensing extraction model for ground object elements corresponding to the target area and the remote sensing image of the target area, a remote sensing extraction result of the ground object elements in the target area is obtained. When the training sample data for training the source domain model is missing, the source domain model can be domain transferred through the improved domain transfer training, and the features in the remote sensing image of the target area can be transferred to the source domain model, which can significantly improve the adaptability of the source domain model to the target area, thereby improving the accuracy of remote sensing extraction of ground object elements in the target area. Description of the Drawings
[0037] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0038] Figure 1 It is one of the flow diagrams of the domain transfer method for remote sensing extraction of regional surface elements provided by the present invention.
[0039] Figure 2 It is another flow diagram of the domain transfer method for remote sensing extraction of regional surface elements provided by the present invention.
[0040] Figure 3 It is yet another flow diagram of the domain transfer method for remote sensing extraction of regional surface elements provided by the present invention.
[0041] Figure 4 It is the framework diagram of constructing a spatio-temporal consistency loss function in the domain transfer method for remote sensing extraction of regional surface elements provided by the present invention.
[0042] Figure 5 It is one of the effect comparison diagrams between the domain migration method for remote sensing extraction of regional surface elements provided by the present invention and the traditional domain migration method for remote sensing extraction of regional surface elements in related technologies.
[0043] Figure 6 It is the second of the effect comparison diagrams between the domain migration method for remote sensing extraction of regional surface elements provided by the present invention and the traditional domain migration method for remote sensing extraction of regional surface elements in related technologies.
[0044] Figure 7 It is a schematic structural diagram of the domain migration device for remote sensing extraction of regional surface elements provided by the present invention.
[0045] Figure 8 It is a schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners
[0046] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.
[0047] In the description of the invention, it should be noted that unless otherwise clearly defined and limited, the terms "installed", "connected" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0048] In the description of the present application, the terms "first", "second", etc. are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are usually of the same kind, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, in the description of the present application, " / " indicates at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.
[0049] It should be noted that the methods for optimizing the trained ground feature extraction model based on domain transfer technology mainly include the following three categories: domain transfer methods based on image transfer, domain transfer methods based on adversarial training, and domain transfer methods based on self-training. Among them, the domain transfer method based on self-training is the most commonly used method in practical applications.
[0050] The domain transfer method based on self-training generally uses the trained model to generate pseudo labels for test data, and uses the pseudo labels to iteratively fine-tune the model, thereby gradually adapting the knowledge of the training sample data to the test data. Therefore, the quality of the pseudo labels is the key to affecting the adaptability of the model.
[0051] However, in the related art, when optimizing the model based on the self-training domain transfer method, the limitation of defining the reliability of pseudo labels only based on the output information of the current model iteration is not considered. Among them, the unreliability of pseudo labels mainly manifests in two aspects, namely, pseudo label inconsistency and knowledge forgetting.
[0052] On the one hand, experimental observations show that the model is prone to produce inconsistent semantic categories under different local spatial perspectives, and the semantic inconsistency generated by pseudo-labels during the iteration process is a key factor leading to unstable model training and even performance degradation.
[0053] On the other hand, knowledge forgetting is another key factor that causes instability in the model optimization process. Since the self-training process lacks effective labeled data, model knowledge becomes the only available supervisory information for the domain transfer process. As the training process continues to iterate, the model will gradually deviate from the knowledge in the training sample data, resulting in the output information of the current model iteration being inaccurate.
[0054] Therefore, the optimization effect of optimizing the trained model based on the traditional domain transfer technology in the related art is not good. Based on the optimized model, it is still difficult to accurately perform remote sensing extraction of ground object elements on the test data when there is a difference in feature distribution between the training sample data and the test data. How to better optimize the trained model so as to more accurately perform remote sensing extraction of ground object elements on the test data when there is a difference in feature distribution between the training sample data and the test data is a technical problem to be solved in this field.
[0055] In response to this, the present invention provides a domain transfer method for remote sensing extraction of ground feature elements in a large area. The remote sensing method for ground feature elements provided by the present invention addresses the problem of inconsistent predictions of the source domain model in different spatial contexts by proposing a spatial multi-perspective consistency learning mechanism. Through the processes of spatial multi-perspective enhancement and prediction fusion, stable pseudo-labels that are unified across multiple perspectives are obtained. To avoid the problem of knowledge forgetting caused by the lack of supervision information guidance during the domain transfer process, a temporal dynamic consistency learning mechanism is proposed, which introduces the model knowledge from historical moments to regularize the training direction of the model, enabling the model to have the ability to self-correct the noise of the current pseudo-labels. At the same time, to alleviate the cognitive bias problem caused by domain shift in the classification task of large-area surface element extraction, a spatio-temporal contrast strategy is designed, which uses contrastive learning to constrain the internal semantic associations of different category features, improving the category discrimination representation of the remote sensing extraction classification model for surface elements in the target domain.
[0056] The following combines Figures 1-6 to describe the domain transfer method for remote sensing extraction of regional surface feature elements provided by the present invention.
[0057] Figure 1 is one of the schematic flowcharts of the domain transfer method for remote sensing extraction of regional surface feature elements provided by the present invention. As Figure 1 shown, the method includes the following: Step 101, obtain the remote sensing image of the target area as the target domain image.
[0058] It should be noted that the execution subject of the embodiments of the present invention is a domain transfer device for remote sensing extraction of regional surface feature elements. The above-mentioned domain transfer device for remote sensing extraction of regional surface feature elements can be configured in electronic devices such as computers or servers.
[0059] Specifically, the target area is the remote sensing extraction object of the domain transfer method for remote sensing extraction of regional surface feature elements provided by the present invention. Based on the remote sensing image of the target area, the domain transfer method for remote sensing extraction of regional surface feature elements provided by the present invention can obtain information such as the type, quantity, and spatial distribution of the ground feature elements in the target area.
[0060] It should be noted that the ground feature elements in the embodiments of the present invention may include, but are not limited to, water systems, residential areas and facilities, transportation facilities, pipeline facilities, boundaries, landforms, vegetation, and soil quality, etc. The water system may include, but is not limited to, natural water bodies such as rivers, lakes, and reservoirs and their affiliated facilities, such as dams, bridges, etc. The residential areas and facilities may include, but are not limited to, cities, villages, buildings, and public facilities, etc. The transportation facilities may include, but are not limited to, transportation facilities such as roads, railways, bridges, and tunnels and their affiliated facilities. The pipeline facilities may include, but are not limited to, various pipelines and lines, such as oil pipelines, power transmission lines, etc. The boundaries may include, but are not limited to, administrative boundaries such as national boundaries, provincial boundaries, and county boundaries. The vegetation and soil quality may include, but are not limited to, vegetation-covered areas such as forests and grasslands and different types of soil.
[0061] It can be understood that the target area in the embodiments of the present invention can be determined based on actual needs. There is no specific limitation on the target area in the embodiments of the present invention.
[0062] It should be noted that the target area in the embodiments of the present invention can be a region with a relatively large area. For example, the target area can be a region under the jurisdiction of a certain administrative unit.
[0063] It should be noted that the remote sensing image of the target area in the embodiments of the present invention can be collected by a remote sensing satellite or by an image sensor installed on an unmanned aerial vehicle. There is no limitation on the specific acquisition method of the remote sensing image of the target area in the embodiments of the present invention.
[0064] In the embodiments of the present invention, the remote sensing image of the target area can be obtained in various ways. For example, the remote sensing image of the target area can be obtained based on the input of the user; or, the remote sensing image of the target area sent by other electronic devices can also be received. There is no limitation on the specific method of obtaining the remote sensing image of the target area in the embodiments of the present invention.
[0065] After obtaining the remote sensing image of the target area, the remote sensing image of the target area can be determined as the target domain image data.
[0066] Step 102: Based on the target domain image, obtain multiple reconstructed sample images with different local spatial perspectives of the target area.
[0067] Specifically, after obtaining the target domain image, multiple reconstructed sample images with different local spatial perspectives can be obtained based on the target domain image through methods such as image preprocessing and deep learning technology.
[0068] It should be noted that different local spatial perspectives in the embodiments of the present invention can be understood as at least one difference in the spatial position, observation angle, and observation scale when observing a local area of the target region. Correspondingly, multiple reconstructed sample images with different local spatial perspectives can include images of multiple different local areas in the target region, and at least one of the shooting positions, shooting angles, and shooting scales of the images of the above-mentioned local areas is different.
[0069] As an optional embodiment, based on the target domain image, multiple reconstructed sample images with different local spatial perspectives of the target region are obtained, including: cropping the target domain image based on a sliding window of a preset size and a predefined overlap rate to obtain multiple original sample images.
[0070] Specifically, the sliding window in the embodiments of the present invention is rectangular, and the length and width of the above sliding window are preset based on prior knowledge and / or actual situations. In the embodiments of the present invention, the size ( ) of the above sliding window is not specifically limited.
[0071] The overlap rate in the embodiments of the present invention refers to the overlap ratio of the above sliding window before and after sliding. The specific value of the overlap rate in the embodiments of the present invention can be predefined based on prior knowledge and / or actual situations. In the embodiments of the present invention, the specific value of the overlap rate is not limited.
[0072] Based on the above preset size and predefined overlap rate, the sliding step of the above sliding window can be calculated, and then starting from the upper left corner of the target domain image, the above sliding window can be slid according to the above sliding step, and the position where the above sliding window is located after each sliding is determined as the position of each original sample image in the target domain image.
[0073] After determining the position of each original sample image in the target domain image, the target domain image can be cropped to obtain each original sample image. Each original sample image can be denoted as . Among them, represents the rd original sample image, represents a positive integer not greater than ; represents the total number of each original sample image; represents the rd original sample image with a size of and 3 channels.
[0074] For each original sample image, multiple sub-region bounding boxes with the same size, located in different regions of each original sample image and overlapping with each other, are randomly generated in each original sample image.
[0075] Specifically, for each original sample image, it can be denoted as the th original sample image in which randomly generated sub-region bounding boxes with the same size, located in different regions of the th original sample image and overlapping with each other. Among them, represents a positive integer greater than 2, and the value of can be determined based on prior knowledge and / or actual situations. For example,
[0076] the th sub-region bounding box in the th original sample image can be denoted as where represents a positive integer not greater than ; the sub-region bounding boxes in the th original sample image can be denoted as
[0077] Based on the sub-region bounding boxes in each original sample image and the bounding box of each original sample image, using the Region of Interest (ROI) Align technology, each reconstructed sample image corresponding to each original sample image is obtained.
[0078] It should be noted that the ROI Align technology is a technology widely used in the fields of object detection and image recognition, mainly used to solve the problem of position mismatch caused by quantization operations in the Region of Interest Pooling (ROI Pooling) operation. The core idea of the ROI Align technology is to cancel the quantization operation and use the method of bilinear interpolation to obtain the image values at pixel points with floating-point coordinates, thereby transforming the entire feature aggregation process into a continuous operation.
[0079] The th sub-region bounding boxes in the original sample image and the th The boundary box of the original sample image After being input into the ROI Align network, the original sample image corresponding reconstructed sample images can be obtained. Among them, the original sample image corresponding to each reconstructed sample image has a size of , and the number of original sample images corresponding to each reconstructed sample image is sheets.
[0080] The original sample image corresponding to the th reconstructed sample image can be expressed as , where represents a positive integer not greater than . The original sample images corresponding to each reconstructed sample image can be expressed as
[0081] Step 103: Based on each reconstructed sample image and the teacher network model, perform improved domain transfer training on the student network model to obtain a trained student network model as the remote sensing extraction model of the ground feature elements corresponding to the target area. The model structures of the teacher network model and the student network model are the same as those of the source domain model, and the initial model parameters of the teacher network model and the student network model are the model parameters of the source domain model. The source domain model is trained based on the remote sensing images of the sample area and the remote sensing images of the labeled sample area. The labeling information in the remote sensing images of the labeled sample area is used to indicate at least one of the type, identification, quantity, and regional range of the ground feature elements.
[0082] Specifically, after obtaining the remote sensing images of the sample area and the remote sensing images of the labeled sample area, based on a sliding window with a size of and an overlap rate of , the sliding step length of the sliding window can be calculated. Then, starting from the upper left corner of the remote sensing image of the sample area and the remote sensing image of the labeled sample area respectively, slide the sliding window according to the sliding step length, and determine the position of each sliding window after each slide in the remote sensing image of the sample area as the position of each source domain sample image in the remote sensing image of the sample area and the position of each labeled source domain sample image in the remote sensing image of the labeled sample area.
[0083] After determining the positions of each source domain sample image in the remote sensing image of the sample area and the positions of each labeled source domain sample image in the remote sensing image of the labeled sample area, the remote sensing image of the sample area and the remote sensing image of the labeled sample area can be cropped respectively to obtain each source domain sample image. Each source domain sample image and each labeled source domain sample image can be denoted as . Among them, represents the th source domain sample image, represents the th labeled source domain sample image, represents a positive integer not greater than ; represents the total number of source domain sample images; represents the th source domain sample image with a size of , and the number of channels is 3; represents the th labeled source domain sample image with a size of , and the number of channels is 3.
[0084] Using the source domain sample images as training sample data and the labeled source domain sample images as sample labels to perform supervised training on the initial model, and the training loss is pixel-level cross-entropy loss , a trained source domain model can be obtained. Among them, represents the model parameters of the trained source domain model.
[0085] It should be noted that the initial model in the embodiments of the present invention can be constructed based on a traditional deep learning network model. For example, the above initial model can be constructed based on the SegFormer model (a deep learning model based on the Transformer architecture).
[0086] The initial model may include a feature extractor and a classifier. The feature extractor can be used to extract the features of the input remote sensing image to obtain the corresponding feature image of the remote sensing image. The classifier can be used to identify and classify the ground object elements based on the above feature image, so as to output the remote sensing recognition result of the ground object elements of the input image.
[0087] It can be understood that the sample area in the embodiments of the present invention may be the same as or different from the target area. However, the remote sensing images of the sample area and the target area are remote sensing images with different time phases.
[0088] In the embodiments of the present invention, the trained source domain model can be respectively determined as the teacher network model and the student network model , the model parameters of the trained source domain model are respectively determined as the initial values of the model parameters of the teacher network model and the initial values of the model parameters of the student network model and the student network model of the model parameters initial value .
[0089] Figure 2 is the second schematic flow chart of the domain transfer method for remote sensing extraction of regional surface elements provided by the present invention. As Figure 2 shown, as an optional embodiment, based on each reconstructed sample image and the teacher network model, improved domain transfer training is performed on the student network model to obtain a trained student network model as a remote sensing extraction model for the ground object elements corresponding to the target area, including: in the k th training, based on each reconstructed sample image, using the teacher network model and the student network model in the k th training, obtain each pseudo-label corresponding to each original sample image in the k th training.
[0090] Based on each reconstructed sample image, using the teacher network model and the student network model in the k th training, obtain each pseudo-label corresponding to each original sample image in the k th training and the embedded features corresponding to each reconstructed sample image and the embedded features corresponding to each original sample image for each original sample image in the k th training.
[0091] Based on each pseudo-label corresponding to each original sample image in the k th training, construct the spatio-temporal consistency loss function corresponding to the k th training, and based on the embedded features corresponding to each reconstructed sample image and the embedded features corresponding to each original sample image for each original sample image in the k th training, construct the spatio-temporal contrast loss function corresponding to the k th training.
[0092] Figure 3 is the third schematic flow chart of the domain transfer method for remote sensing extraction of regional surface elements provided by the present invention. Figure 4 is the framework diagram for constructing the spatio-temporal consistency loss function in the domain transfer method for remote sensing extraction of regional surface elements provided by the present invention. As Figure 3 and Figure 4 shown, as an optional embodiment, based on each reconstructed sample image, using the kThe teacher network model and the student network model in the k th training are obtained, and each pseudo-label corresponding to each original sample image in the k th training is obtained, including: in the
[0093] th training, based on the exponential moving average method, the model parameters of the teacher network model are updated.
[0094]
[0095] It should be noted that the student network model in the embodiments of the present invention is a trainable network, and the teacher network model updates the model parameters by using the dynamic update mechanism of Dynamic Teacher-Student Teacher (DT-ST). Specifically, the exponential moving average (EMA) method is used for the update method. The mathematical expression for the teacher network model to update the model parameters is as follows: represents the smoothing coefficient (usually close to 1, such as 0.99), which is used to control the update speed; represents the model parameters of the teacher network model in the k th training; represents the model parameters of the student network model in the k th training; represents the model parameters of the teacher network model in the k-m th training; m represents a predefined integer value, for example m can take a value of 3.
[0096] It should be noted that in the case of , k-m takes a value of k .
[0097] Each reconstructed sample image corresponding to each original sample image is respectively input into the teacher network model and the student network model, and the first predicted image of each reconstructed sample image corresponding to each original sample image in the k th training output by the teacher network model and the second predicted image of each reconstructed sample image corresponding to each original sample image in the k th training output by the student network model are obtained.
[0098] Specifically, for the th original sample image corresponding to the th reconstructed sample image , in the th training, the th original sample image The corresponding reconstructed sample image After being input into the teacher network model and the student network model respectively, the k th training original sample image The corresponding reconstructed sample image The first predicted image and the k th training original sample image The corresponding reconstructed sample image second predicted image can be obtained respectively.
[0099] It should be noted that the original sample image The corresponding reconstructed sample image The first predicted image can be the original sample image after annotation The corresponding reconstructed sample image . The original sample image The corresponding reconstructed sample image The second predicted image can also be the original sample image after annotation The corresponding reconstructed sample image .
[0100] Using the reverse process of the ROI Align technology, data fusion is performed on the first predicted images of the reconstructed sample images corresponding to each original sample image in the k th training to obtain the fusion predicted images corresponding to each original sample image in the k th training.
[0101] It should be noted that the reverse process of the ROI Align technology is usually called Inverse ROI Align or ROI Unalign, and its goal is to remap the feature map processed by the ROI Align technology back to the original image space.
[0102] The th training original sample image After the first predicted image of each reconstructed sample image is input into the reverse ROI Align network, the reverse ROI Align network can k In the training Original sample images The corresponding reconstructed sample images are fused and the k In the training Original sample images The overlapping areas in the corresponding reconstructed sample images are fused by predictive average, and then the size of the output of the above reverse ROI Align network can be obtained as No. k In the training Original sample images The corresponding fused predicted image.
[0103] The mathematical expression of the calculation process of the reverse ROI Align network is as follows:
[0104]
[0105] in, Indicated in In the first training session, Original sample images The corresponding Reconstructed sample image After inputting the teacher network model, the teacher network model outputs k In the training Original sample images The corresponding Reconstructed sample image The first predicted image of represents the mask function, which is obtained based on the sub-region bounding box; Indicates k In the training Original sample images The corresponding fused predicted image.
[0106] Based on k The fused prediction image corresponding to each original sample image in the training is obtained by using the ROI Align technology. k Each secondary reconstructed image corresponding to each original sample image in the training is used as the k Each pseudo label corresponding to each original sample image in the training.
[0107] Specifically, obtain the k In the training Original sample images Corresponding fusion prediction image After that, the k th training, the th original sample image corresponding fused prediction image is input into the ROI Align network, and the k th training, the th original sample image corresponding each second reconstruction image can be obtained.
[0108] It should be noted that in the embodiment of the present invention, the k th training, the th original sample image corresponding size of the second reconstruction image is the same as that of the k th training, the th original sample image corresponding th reconstructed sample image first prediction image.
[0109] After obtaining the k th training, the th original sample image corresponding each second reconstruction image, the k th training, the th original sample image corresponding each second reconstruction image can be respectively determined as the k th training, the th original sample image corresponding each pseudo-label.
[0110] In the embodiment of the present invention, after reconstructing multiple reconstructed sample images with different local spatial perspectives of the target region based on the ROI Align technology, the above-mentioned reconstructed sample images are respectively input into the student network model and the teacher network model. The output of the teacher network model undergoes prediction average fusion and label re-alignment to obtain pseudo-labels with consistent semantics under different spatial perspectives, so as to perform spatial multi-perspective consistency training on the student network model.
[0111] As an optional embodiment, based on each pseudo-label corresponding to each original sample image in the th training, a spatio-temporal consistency loss function corresponding to the k th training is constructed, including: in the k th training, each original sample image is input into the last target historical teacher network model to obtain thek The third predicted image corresponding to each original sample image in the
[0112] training. Among them, the last target historical teacher network model is the target historical teacher network model arranged at the end in the historical teacher network model queue.
[0113] The historical teacher network model queue is obtained based on the following steps: When it is determined that the first training ends, the teacher network model in the first training is determined as the target historical teacher network model corresponding to the first training, and the target historical teacher network model corresponding to the first training is added to the historical teacher network model queue.
[0114] When it is determined that the th training ends, and it is determined that the difference between the training times corresponding to the target historical teacher network model arranged at the head in the historical teacher network model queue is a preset value, the teacher network model in the th training is determined as the target historical teacher network model. .
[0115] When the number of target historical teacher networks in the historical teacher network model queue is less than the number threshold, the target historical teacher network model corresponding to the th training is inserted into the head of the historical teacher network model queue. When the number of target historical teacher networks in the historical teacher network model queue is equal to the number threshold, the target historical teacher network model corresponding to the th training is inserted into the head of the historical teacher network model queue, and at the same time, the target historical teacher network model arranged at the end in the historical teacher network model queue is removed.
[0116] Specifically, the number threshold and the preset value in the embodiments of the present invention can be determined based on prior knowledge and / or actual situations. In the embodiments of the present invention, the specific values of the number threshold and the preset value are not limited.
[0117] Based on the third predicted image corresponding to each original sample image in the th training and the fused predicted image corresponding to each original sample image in the k th training, the consistency weight map corresponding to each original sample image in the k th training is calculated.
[0118] Specifically, based on the third predicted image corresponding to each original sample image in the k th training and the fused predicted image corresponding to each original sample image in the kFor each first predicted image of each reconstructed sample image corresponding to each original sample image in the k -th training, calculate the -th original sample image
[0119]
[0120] wherein, represents the consistency weight map corresponding to the k -th original sample image in the -th training; k represents the first predicted image of each reconstructed sample image corresponding to the -th original sample image in the -th training; represents the third predicted image corresponding to the -th original sample image output by the end target historical teacher network model;
[0121] represents the -th -th training; k Based on the consistency weight map corresponding to each original sample image in the
[0122] -th training and each pseudo-label corresponding to each original sample image in the k -th training, construct the spatio-temporal consistency loss function corresponding to the -th training. Specifically, based on the consistency weight map corresponding to the -th k -th original sample image in the k -th training, weights can be assigned to each pseudo-label corresponding to the -th original sample image k in the -th training, and then based on each pseudo-label corresponding to the
[0123]
[0124] wherein, Denote the k weight value of the point with coordinate in the corresponding consistency weight map of the th original sample image in the th training; Denote the pixel value of the point with coordinate in the corresponding fused prediction image of the th original sample image in the th training; Denote that in the th training, after inputting the th reconstructed sample image corresponding to the th original sample image into the student network model, the th reconstructed sample image output by the student network model, and the second prediction image of the th original sample image corresponding to the th original sample image in the th training; Denote the Softmax function.
[0125]
[0125]
[0125] In the embodiment of the present invention, by maintaining a queue of historical teacher network models, a consistency evaluation of the output predictions of the current teacher network model in training and the last target historical teacher network model at the end of the queue is performed, a consistency weight map in different temporal model states is obtained, and weights are re-assigned to the pseudo-labels, so as to perform temporal dynamic consistency training on the student network model by constructing a consistency loss.
[0126] As an optional embodiment, based on each reconstructed sample image, using the teacher network model and the student network model in the th training, obtain each pseudo-label corresponding to each original sample image in the th training and the embedding features corresponding to each reconstructed sample image corresponding to each original sample image and the embedding features corresponding to each original sample image in the th training, including: in the th training, after inputting each reconstructed sample image corresponding to each original sample image into the student network model, obtain the th feature image corresponding to each reconstructed sample image corresponding to each original sample image output by the feature extractor in the student network model, and in the th training, after inputting each original sample image into the last target historical teacher network model, obtain the The feature images corresponding to each original sample image in the
[0127] It can be understood that the source domain model in the embodiments of the present invention includes a feature extractor and a classifier, and the network architectures of the student network model and the source domain model are the same. Therefore, both the student network model and the teacher network model in the embodiments of the present invention include a feature extractor and a classifier.
[0128] In the th training, the th original sample image corresponding th reconstructed sample image After being input into the student network model, the feature image corresponding to the th training and the th original sample image corresponding th reconstructed sample image can be obtained.
[0129] In the th training, after the th original sample image is input into the last target historical teacher network model, the feature image corresponding to the k th training and the th original sample image can be obtained.
[0130] The feature images corresponding to each reconstructed sample image and the feature images corresponding to each original sample image in the k th training are respectively projected into the feature space of a preset dimension to obtain the embedded features corresponding to each reconstructed sample image and the embedded features corresponding to each original sample image in the k th training.
[0131] It should be noted that the preset dimension in the embodiments of the present invention can be determined based on prior knowledge and / or actual situations, and the specific value of the above preset dimension in the embodiments of the present invention is not limited.
[0132] Optionally, the preset dimension in the embodiments of the present invention can be 128.
[0133] Obtain the feature images corresponding to each reconstructed sample image corresponding to each original sample image in the k th training and the feature images corresponding to each original sample image, as well as the kAfter the feature images corresponding to each original sample image in the k -th training, a projection head composed of two-layer Multilayer Perceptron (MLP) can be used to project the feature images corresponding to each reconstructed sample image and the feature images corresponding to each original sample image in the k -th training into a 128-dimensional feature space respectively, so as to obtain the embedded features corresponding to each reconstructed sample image and the embedded features corresponding to each original sample image in the k -th training respectively and the embedded features corresponding to each original sample image .
[0134] It should be noted that the encoded embedding is the feature map output by the encoder (feature extractor) of the model. The output of the model contains semantic categories. After resampling the feature map and the output of the model to the same size, the feature regions with the same semantic category and the feature regions with different semantic categories can be obtained, so that positive samples, negative samples and anchors can be divided.
[0135] After obtaining the embedded features corresponding to each reconstructed sample image and the embedded features corresponding to each original sample image in the k -th training and the embedded features corresponding to each original sample image , the anchors, positive samples and negative samples can be constructed based on the embedded features corresponding to each reconstructed sample image and the embedded features corresponding to each original sample image in the k -th training and the embedded features corresponding to each original sample image .
[0136] Among them, the anchor is the encoded embedding of the student network model, the positive sample is the encoded embedding with the same semantic category as the anchor, and the encoded embeddings with other different semantic categories are negative samples. The spatio-temporal contrast loss function corresponding to the -th training is constructed, and the specific formula is as follows:
[0137]
[0138] Among them, represents the value of the spatio-temporal contrast loss function corresponding to the -th training; represents the semantic category of the embedded feature corresponding to each reconstructed sample image and each original sample image in the -th training; represents the semantic category of the embedded feature corresponding to each original sample image in the -th training; represents the number of embedded elements, where the embedded elements refer to the pixel units of the feature map; represents the temperature coefficient; represents a positive integer; represents the th embedded vector corresponding to the
[0139] In the embodiment of the present invention, the embedded features are obtained by using the feature extractor of the student network model, and the feature image is obtained by using the last target historical teacher network model. The embedded features of the student network model are used as anchor samples, and the encoded embeddings projected by the pseudo-labels and having the same semantic category as the anchor samples are used as positive samples, and the encoded embeddings belonging to other semantic categories are used as negative samples. A spatio-temporal contrast loss is constructed to perform spatio-temporal contrast training on the student network model.
[0140] Based on the spatio-temporal consistency loss function and the spatio-temporal contrast loss function corresponding to the th training, combined with the minimum entropy loss function, the target loss function corresponding to the th training is constructed.
[0141] Specifically, the formula of the minimum entropy loss function corresponding to the th training is expressed as follows:
[0142]
[0143] where, represents the th semantic category; represents the sum of the semantic categories.
[0144] The formula of the target loss function corresponding to the th training is expressed as follows:
[0145]
[0146] where, represents the value of the target loss function corresponding to the th training; and are set as the relative contribution degrees of the spatio-temporal contrast loss function and the minimum entropy loss function in the model gradient update optimization process.
[0147] Calculate the function value of the target loss function corresponding to the th training. Based on the function value of the target loss function corresponding to the th training, determine that the student network model has not converged or is not greater than the maximum number of training times, and update the model parameters of the student network model. Increment by 1, and return to execute at the k th training, based on each reconstructed sample image, using the teacher network model and the student network model in the k th training, obtaining each pseudo-label corresponding to each original sample image in the k th training and the embedded features corresponding to each reconstructed sample image and the embedded features corresponding to each original sample image in the k th training. In the step of determining that the student network model converges or is greater than the maximum number of training times based on the function value of the target loss function corresponding to the th training, it is determined that the student network model has been trained well, and the trained student network model is determined as the remote sensing extraction model of the ground object elements corresponding to the target area.
[0148] Step 104: Based on the remote sensing extraction model of the ground object elements corresponding to the target area and the remote sensing image of the target area, obtain the remote sensing extraction result of the ground object elements in the target area.
[0149] As an optional embodiment, based on the remote sensing extraction model of the ground object elements corresponding to the target area and the remote sensing image of the target area, obtaining the remote sensing extraction result of the ground object elements in the target area includes: inputting each original sample image into the remote sensing extraction model of the ground object elements corresponding to the target area, and obtaining the predicted image corresponding to each original sample image output by the remote sensing extraction model of the ground object elements corresponding to the target area.
[0150] Stitch the predicted images corresponding to the original sample images to obtain the predicted image of the target area.
[0151] Extract information from the predicted image of the target area to obtain the remote sensing extraction result of the ground object elements in the target area.
[0152] In the embodiment of the present invention, by obtaining the remote sensing image of the target area, after using it as the target domain image, based on the target domain image, multiple reconstructed sample images with different local spatial perspectives of the target area are obtained. Based on each reconstructed sample image and the teacher network model, the domain transfer training of the student network model is improved, and the trained student network model is obtained as the remote sensing extraction model of the ground object elements corresponding to the target area. Furthermore, based on the remote sensing extraction model of the ground object elements corresponding to the target area and the remote sensing image of the target area, the remote sensing extraction result of the ground object elements in the target area is obtained. It can, in the case of the lack of training sample data for training the source domain model, through the improved domain transfer training, perform domain transfer on the source domain model, transfer the features in the remote sensing image of the target area to the source domain model, and can significantly improve the adaptability of the source domain model to the target area, thereby improving the accuracy of remote sensing extraction of ground object elements in the target area.
[0153] The flow chart of the domain transfer method for remote sensing extraction of regional surface elements provided by the present invention is different from the traditional domain transfer technology in the related art, which only considers the reliability of pseudo-labels defined by the output information of the current model iteration. The improved domain transfer method in the present invention includes two parts: spatial multi-perspective consistency learning and temporal dynamic consistency learning. Spatial multi-perspective consistency learning improves the context consistency and reliability of pseudo-labels through the fusion alignment of multi-perspective enhancement and output prediction. Temporal dynamic consistency learning regularizes the evolution direction of the model by constraining the consistency between the current and historical model knowledge, preventing the model from gradually forgetting knowledge due to the lack of supervision information.
[0154] Moreover, the domain transfer method for remote sensing extraction of regional surface elements provided by the present invention designs a new spatio-temporal contrast strategy. Different from the traditional domain transfer technology in the related art that lacks semantic association constraints at the spatio-temporal level, the spatio-temporal contrast strategy in the present invention constrains the intrinsic semantic association of different category features through model embeddings at different times and different spatial perspectives, so as to improve the category discrimination representation of the remote sensing extraction classification model of surface elements in the target domain, and further improve the generalization ability of the model.
[0155] Compared with the related art, the domain transfer method for remote sensing extraction of regional surface elements provided by the present invention solves the problem of inconsistent predictions of the remote sensing extraction classification model of surface elements in different spatial contexts through the spatial multi-perspective consistency learning mechanism. Through the spatial multi-perspective enhancement and prediction fusion process, more stable and reliable multi-perspective unified pseudo-labels can be obtained, making the domain transfer model more stable and adaptable; through the temporal dynamic consistency learning mechanism, the problem of knowledge forgetting caused by the lack of supervision information guidance in the domain transfer process is solved. The model knowledge of historical moments is introduced to regularize the training direction of the model, enabling the model to have the self-correction ability for the noise of current pseudo-labels, making the element extraction classification results of the model in the target domain more complete and the false detection and misdetection rates lower; through the spatio-temporal contrast learning strategy, the cognitive bias problem caused by domain shift in the large-area surface element extraction classification task is solved. By using contrast learning to constrain the intrinsic semantic association of different category features, the model has stronger discrimination ability between different surface elements with small feature differences.
[0156] Table 1 is one of the accuracy comparison tables between the domain transfer method for remote sensing extraction of regional surface elements provided by the present invention and the traditional domain transfer method in the related art. Figure 5 It is one of the effect comparison diagrams between the domain transfer method for remote sensing extraction of regional surface elements provided by the present invention and the traditional domain transfer method for remote sensing extraction of regional surface elements in the related art.
[0157] The domain adaptation method for remote sensing extraction of regional surface elements provided by the present invention is applied to the loveDA urban-rural domain adaptation dataset, and the application effect is compared with traditional domain adaptation methods such as HCL, IAPC, and DT-ST. The accuracy evaluation results and effect evaluation results are shown in Table 1 and Figure 5 as shown below.
[0158] Table 1 Comparison table of accuracy between the domain adaptation method for remote sensing extraction of regional surface elements provided by the present invention and traditional domain adaptation methods in related technologies (Part 1)
[0159]
[0160] As shown in Table 1 and Figure 5 as shown below, compared with the traditional domain adaptation methods in related technologies, the accuracy of the domain adaptation method for remote sensing extraction of regional surface elements provided by the present invention in remote sensing extraction of surface elements is significantly improved.
[0161] Table 2 is the comparison table of accuracy between the domain adaptation method for remote sensing extraction of regional surface elements provided by the present invention and traditional domain adaptation methods in related technologies (Part 2). Figure 6 Figure 2 is the comparison chart of the effects between the domain adaptation method for remote sensing extraction of regional surface elements provided by the present invention and traditional domain adaptation methods for remote sensing extraction of regional surface elements in related technologies (Part 2).
[0162] Furthermore, the domain adaptation method for remote sensing extraction of regional surface elements provided by the present invention is applied to the large-area and large-scale surface element extraction and classification dataset, and the application effect is compared with traditional domain adaptation methods such as HCL, IAPC, and DT-ST. The accuracy evaluation results and effect evaluation results are shown in Table 2 and Figure 6 as shown below.
[0163] Table 2 Comparison table of accuracy between the domain adaptation method for remote sensing extraction of regional surface elements provided by the present invention and traditional domain adaptation methods in related technologies (Part 2)
[0164]
[0165] As shown in Table 2 and Figure 6 as shown below, compared with the traditional domain adaptation methods in related technologies, the accuracy of the domain adaptation method for remote sensing extraction of regional surface elements provided by the present invention in remote sensing extraction of surface elements is significantly improved.
[0166] Figure 7 Figure 3 is the structural schematic diagram of the domain adaptation device for remote sensing extraction of regional surface elements provided by the present invention. The following combines Figure 7The domain transfer device for remote sensing extraction of regional surface elements provided by the present invention is described. The domain transfer device for remote sensing extraction of regional surface elements described below can be correspondingly referred to the domain transfer method for remote sensing extraction of regional surface elements provided by the present invention described above. As Figure 7 shown, the device includes: a data acquisition module 701, a sample reconstruction module 702, a model training module 703, and a remote sensing extraction module 704.
[0167] The data acquisition module 701 is configured to acquire a remote sensing image of a target area as a target domain image.
[0168] The sample reconstruction module 702 is configured to obtain multiple reconstructed sample images with different local spatial perspectives of the target area based on the target domain image.
[0169] The model training module 703 is configured to perform improved domain transfer training on the student network model based on each reconstructed sample image and the teacher network model to obtain a trained student network model as the remote sensing extraction model of the ground feature corresponding to the target area. The model structures of the teacher network model and the student network model are the same as the model structure of the source domain model, and the initial model parameters of the teacher network model and the student network model are the model parameters of the source domain model. The source domain model is trained based on the remote sensing image of the sample area and the remote sensing image of the labeled sample area. The labeling information in the remote sensing image of the labeled sample area is used to indicate at least one of the type, identifier, quantity, and area range of the ground feature.
[0170] The remote sensing extraction module 704 is configured to obtain the remote sensing extraction result of the ground feature of the target area based on the remote sensing extraction model of the ground feature corresponding to the target area and the remote sensing image of the target area.
[0171] Specifically, the data acquisition module 701, the sample reconstruction module 702, the model training module 703, and the remote sensing extraction module 704 are electrically connected.
[0172] In the domain transfer device for remote sensing extraction of regional surface elements in the embodiments of the present invention, after obtaining the remote sensing image of the target area as the target domain image, based on the target domain image, multiple reconstructed sample images with different local spatial perspectives of the target area are obtained. Based on each reconstructed sample image and the teacher network model, an improved domain transfer training is performed on the student network model to obtain a trained student network model as the remote sensing extraction model of the ground object elements corresponding to the target area. Then, based on the remote sensing extraction model of the ground object elements corresponding to the target area and the remote sensing image of the target area, the remote sensing extraction result of the ground object elements in the target area is obtained. When the training sample data for training the source domain model is missing, through the improved domain transfer training, the domain of the source domain model can be transferred, and the features in the remote sensing image of the target area can be transferred to the source domain model, which can significantly improve the adaptability of the source domain model to the target area, thereby improving the accuracy of remote sensing extraction of ground object elements in the target area.
[0173] Figure 8 An example of the physical structure diagram of an electronic device is shown as Figure 8 shown. The electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communication interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 can call the logical instructions in the memory 830 to execute the domain transfer method for remote sensing extraction of regional surface elements. The method includes: obtaining the remote sensing image of the target area as the target domain image; based on the target domain image, obtaining multiple reconstructed sample images with different local spatial perspectives of the target area; based on each reconstructed sample image and the teacher network model, performing improved domain transfer training on the student network model to obtain a trained student network model as the remote sensing extraction model of the ground object elements corresponding to the target area. The model structures of the teacher network model and the student network model are the same as the model structure of the source domain model, and the initial model parameters of the teacher network model and the student network model are the model parameters of the source domain model. The source domain model is trained based on the remote sensing image of the sample area and the remote sensing image of the labeled sample area. The labeling information in the remote sensing image of the labeled sample area is used to indicate at least one of the type, identifier, quantity, and regional range of the ground object elements; based on the remote sensing extraction model of the ground object elements corresponding to the target area and the remote sensing image of the target area, obtaining the remote sensing extraction result of the ground object elements in the target area.
[0174] In addition, when the logical instructions in the above-mentioned memory 830 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0175] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the domain migration method for remote sensing extraction of regional surface elements provided by the above-mentioned various methods. The method includes: obtaining a remote sensing image of a target area as a target domain image; based on the target domain image, obtaining multiple reconstructed sample images with different local spatial perspectives of the target area; based on each reconstructed sample image and a teacher network model, performing improved domain migration training on a student network model to obtain a trained student network model as a remote sensing extraction model of ground object elements corresponding to the target area. The model structures of the teacher network model and the student network model are the same as the model structure of the source domain model, and the initial model parameters of the teacher network model and the student network model are the model parameters of the source domain model. The source domain model is trained based on the remote sensing image of a sample area and the remote sensing image of the sample area after annotation. The annotation information in the remote sensing image of the sample area after annotation is used to indicate at least one of the type, identifier, quantity, and regional range of the ground object elements; based on the remote sensing extraction model of the ground object elements corresponding to the target area and the remote sensing image of the target area, obtaining the remote sensing extraction result of the ground object elements in the target area.
[0176] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements a domain migration method for remote sensing extraction of regional surface elements provided by the above-mentioned various methods. The method includes: obtaining a remote sensing image of a target area as a target domain image; obtaining multiple reconstructed sample images with different local spatial perspectives of the target area based on the target domain image; performing improved domain migration training on a student network model based on each reconstructed sample image and a teacher network model to obtain a trained student network model as a remote sensing extraction model of ground object elements corresponding to the target area. The model structures of the teacher network model and the student network model are the same as the model structure of the source domain model, and the initial model parameters of the teacher network model and the student network model are the model parameters of the source domain model. The source domain model is trained based on the remote sensing image of a sample area and the remote sensing image of the labeled sample area, and the labeling information in the remote sensing image of the labeled sample area is used to indicate at least one of the type, identifier, quantity, and regional range of the ground object elements; obtaining a remote sensing extraction result of the ground object elements of the target area based on the remote sensing extraction model of the ground object elements corresponding to the target area and the remote sensing image of the target area.
[0177] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.
[0178] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0179] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A domain transfer method for remote sensing extraction of regional surface elements, characterized in that, Including: Obtain a remote sensing image of a target area as a target domain image; Based on the target domain image, obtain multiple reconstructed sample images with different local spatial perspectives of the target area; Based on each of the reconstructed sample images and a teacher network model, perform improved domain transfer training on a student network model to obtain a trained student network model as a remote sensing extraction model of ground object elements corresponding to the target area. The model structures of the teacher network model and the student network model are the same as the model structure of a source domain model, and the initial model parameters of the teacher network model and the student network model are the model parameters of the source domain model. The source domain model is trained based on a remote sensing image of a sample area and a remote sensing image of the annotated sample area, and the annotation information in the remote sensing image of the annotated sample area is used to indicate at least one of the type, identification, quantity, and regional range of the ground object elements; Based on the remote sensing extraction model of ground object elements corresponding to the target area and the remote sensing image of the target area, obtain a remote sensing extraction result of the ground object elements of the target area; The obtaining multiple reconstructed sample images with different local spatial perspectives of the target area based on the target domain image includes: Crop the target domain image based on a sliding window of a preset size and a predefined overlap rate to obtain multiple original sample images; For each original sample image, randomly generate multiple sub-region bounding boxes with the same size, located in different regions of the each original sample image and overlapping with each other; Based on each of the sub-region bounding boxes in the each original sample image and the bounding box of the each original sample image, use the Region of Interest (ROI) Align technique to obtain each reconstructed sample image corresponding to the each original sample image; The performing improved domain transfer training on a student network model based on each of the reconstructed sample images and a teacher network model to obtain a trained student network model as a remote sensing extraction model of ground object elements corresponding to the target area includes: In the k th training, based on each of the reconstructed sample images, using the teacher network model and the student network model in the k th training, obtain each pseudo-label corresponding to each original sample image in the k th training, as well as the embedding features corresponding to each reconstructed sample image and the embedding features corresponding to each original sample image corresponding to each original sample image in the k th training; Based on each pseudo-label corresponding to each original sample image in the k th training, construct the spatio-temporal consistency loss function corresponding to the k th training. Based on the embedding features corresponding to each reconstructed sample image and the embedding features corresponding to each original sample image in the k th training, construct the spatio-temporal contrast loss function corresponding to the k th training; Based on the spatio-temporal consistency loss function and the spatio-temporal contrast loss function corresponding to the k -th training, combined with the minimum entropy loss function, construct the target loss function corresponding to the k -th training; Calculate the function value of the target loss function corresponding to the k th training. When it is determined that the student network model has not converged or k is not greater than the maximum number of training times based on the function value of the target loss function corresponding to the k th training, update the model parameters of the student network model, k increase by 1, and return to execute the step of, in the k th training, based on each of the reconstructed sample images, using the teacher network model and the student network model in the k th training to obtain each pseudo-label corresponding to each original sample image in the k th training, and the embedding features corresponding to each reconstructed sample image and the embedding features corresponding to each original sample image corresponding to each original sample image in the k th training. When it is determined that the student network model has converged or k is greater than the maximum number of training times based on the function value of the target loss function corresponding to the k th training, determine that the student network model has been trained well, and determine the trained student network model as the remote sensing extraction model of the ground feature elements corresponding to the target area.
2. The domain transfer method for remote sensing extraction of regional surface elements according to claim 1, characterized in that Based on each of the reconstructed sample images, using the teacher network model and the student network model in the k -th training, obtain each pseudo-label corresponding to each original sample image in the k -th training, including: In the k th training, based on the exponential moving average method, update the model parameters of the teacher network model; Input each of the reconstructed sample images corresponding to each of the original sample images into the teacher network model and the student network model respectively, and obtain the first predicted image of each of the reconstructed sample images corresponding to each of the original sample images in the k th training by the teacher network model and the second predicted image of each of the reconstructed sample images corresponding to each of the original sample images in the k th training by the student network model; Using the reverse process of the ROI Align technology, for each first predicted image of the reconstructed sample images corresponding to each original sample image described in the k -th training, data fusion is performed to obtain a fused predicted image corresponding to each original sample image described in the k -th training; Based on the fused prediction images corresponding to each original sample image in the k -th training, use the ROI Align technique to obtain each second reconstruction image corresponding to each original sample image in the k -th training, and use it as each pseudo-label corresponding to each original sample image in the k -th training.
3. The domain adaptation method for remote sensing extraction of regional surface elements according to claim 1, characterized in that Based on each pseudo-label corresponding to each original sample image in the k -th training, construct the spatio-temporal consistency loss function corresponding to the k -th training, including: In the k th training, each of the original sample images is input into the last target historical teacher network model to obtain the third predicted image corresponding to each of the original sample images in the k th training; Based on the third predicted image corresponding to each original sample image in the k th training and the fused predicted image corresponding to each original sample image in the k th training, the consistency weight map corresponding to each original sample image in the k th training is calculated; Based on the consistency weight map corresponding to each original sample image in the k th training and each pseudo-label corresponding to each original sample image in the k th training, a spatio-temporal consistency loss function corresponding to the k th training is constructed; Wherein, the last target historical teacher network model is the target historical teacher network model arranged at the end in a historical teacher network model queue; The historical teacher network model queue is obtained based on the following steps: When it is determined that the first training is completed, determine the teacher network model in the first training as the target historical teacher network model corresponding to the first training, and add the target historical teacher network model corresponding to the first training to the historical teacher network model queue; After determining that the m -th training ends and determining that the difference between the training times corresponding to the target historical teacher network model ranked first in the historical teacher network model queue and the m is a preset value, the teacher network model in the m -th training is determined as the target historical teacher network model, ; When the number of target historical teacher networks in the historical teacher network model queue is less than the number threshold, insert the target historical teacher network model corresponding to the m th training to the head of the historical teacher network model queue. When the number of target historical teacher networks in the historical teacher network model queue is equal to the number threshold, insert the target historical teacher network model corresponding to the m th training to the head of the historical teacher network model queue, and at the same time remove the target historical teacher network model arranged at the end of the historical teacher network model queue.
4. The domain transfer method for remote sensing extraction of regional surface elements according to claim 3, characterized in that Based on each of the reconstructed sample images, using the teacher network model and the student network model in the k th training, obtain each pseudo-label corresponding to each original sample image in the k th training and the embedding features corresponding to each reconstructed sample image and the embedding features corresponding to each original sample image for each original sample image in the k th training, including: In the k th training, after inputting each reconstructed sample image corresponding to each original sample image into the student network model, the feature images corresponding to each reconstructed sample image corresponding to each original sample image output by the feature extractor in the student network model in the k th training are obtained. In the k th training, after inputting each original sample image into the last target historical teacher network model, the feature images corresponding to each original sample image output by the feature extractor in the last target historical teacher network model in the k th training are obtained; Project the feature images corresponding to each reconstructed sample image and the feature images corresponding to each original sample image in the k -th training respectively into a feature space of a preset dimension, and obtain the embedded features corresponding to each reconstructed sample image and the embedded features corresponding to each original sample image in the k -th training.
5. The domain adaptation method for remote sensing extraction of regional surface elements according to any one of claims 1 to 4, characterized in that, The obtaining a remote sensing extraction result of the ground object elements of the target area based on the remote sensing extraction model of ground object elements corresponding to the target area and the remote sensing image of the target area includes: Input each original sample image into the remote sensing extraction model of ground object elements corresponding to the target area to obtain a predicted image corresponding to each original sample image output by the remote sensing extraction model of ground object elements corresponding to the target area; Stitch the predicted images corresponding to the original sample images to obtain a predicted image of the target area; Extract information from the predicted image of the target area to obtain the remote sensing extraction result of the ground object elements in the target area.
6. A domain migration device for remote sensing extraction of regional surface elements, characterized in that, It includes: A data acquisition module for acquiring a remote sensing image of the target area as the target domain image; A sample reconstruction module for obtaining multiple reconstructed sample images with different local spatial perspectives of the target area based on the target domain image; A model training module for performing improved domain transfer training on the student network model based on each of the reconstructed sample images and the teacher network model to obtain a trained student network model as the remote sensing extraction model of the ground object elements corresponding to the target area. The model structures of the teacher network model and the student network model are the same as the model structure of the source domain model, and the initial model parameters of the teacher network model and the student network model are the model parameters of the source domain model. The source domain model is trained based on the remote sensing image of the sample area and the remote sensing image of the labeled sample area, and the labeling information in the remote sensing image of the labeled sample area is used to indicate at least one of the type, identification, quantity, and regional scope of the ground object elements; A remote sensing extraction module for obtaining the remote sensing extraction result of the ground object elements in the target area based on the remote sensing extraction model of the ground object elements corresponding to the target area and the remote sensing image of the target area; The sample reconstruction module obtains multiple reconstructed sample images with different local spatial perspectives of the target area based on the target domain image, including: Cropping the target domain image based on a sliding window of a preset size and a predefined overlap rate to obtain multiple original sample images; For each original sample image, randomly generate multiple sub-region bounding boxes with the same size, located in different regions of each original sample image and overlapping with each other; Based on each sub-region bounding box in each original sample image and the bounding box of each original sample image, use the Region of Interest (ROI) Align technology to obtain each reconstructed sample image corresponding to each original sample image; The model training module performs improved domain transfer training on the student network model based on each of the reconstructed sample images and the teacher network model to obtain a trained student network model as the remote sensing extraction model of the ground object elements corresponding to the target area, including: In the k th training, based on each of the reconstructed sample images, using the teacher network model and the student network model in the k th training, obtain each pseudo-label corresponding to each original sample image in the k th training and the embedding features corresponding to each reconstructed sample image corresponding to each original sample image and the embedding features corresponding to each original sample image in the k th training; Based on each pseudo-label corresponding to each original sample image in the k th training, construct the spatio-temporal consistency loss function corresponding to the k th training. Based on the embedding features of each reconstructed sample image corresponding to each original sample image and the embedding features of each original sample image in the k th training, construct the spatio-temporal contrast loss function corresponding to the k th training; Based on the spatio-temporal consistency loss function and the spatio-temporal contrast loss function corresponding to the k th training, combined with the minimum entropy loss function, construct the target loss function corresponding to the k th training; Calculate the k The function value of the target loss function corresponding to the training is based on the k The function value of the target loss function corresponding to the training time determines whether the student network model has not converged or k When the number of training times is not greater than the maximum, the model parameters of the student network model are updated. k Increase by 1 and return to execute the k In the training, based on the reconstructed sample images, the k The teacher network model and the student network model in the training are obtained. k Each pseudo label corresponding to each original sample image in the training k The steps of embedding features corresponding to each reconstructed sample image corresponding to each original sample image and embedding features corresponding to each original sample image in the first training step are based on the first k The function value of the target loss function corresponding to the training time determines whether the student network model converges or k When the number of training times is greater than the maximum number, it is determined that the student network model has been trained, and the trained student network model is determined as the remote sensing extraction model of the land feature corresponding to the target area.
7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the domain transfer method for remote sensing extraction of regional surface elements as described in any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the domain transfer method for remote sensing extraction of regional surface elements as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Ground feature element segmentation model training method and ground feature element segmentation method and device
CN116977633A
Remote sensing image cross-domain small sample classification method based on pseudo label uncertainty perception
CN117152503A