Method and device for identifying surface feature change, equipment and storage medium

By integrating remote sensing images and a pre-trained model with multi-round question-and-answer data, changes in surface features can be identified. This solves the problem that traditional methods have difficulty accurately identifying subtle surface changes under high-resolution remote sensing images, and achieves more accurate change detection.

CN120689768AActive Publication Date: 2025-09-23BEIJING NORMAL UNIVERSITY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510826640.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-23
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

Traditional remote sensing image analysis methods have difficulty accurately identifying subtle changes in the surface in high-resolution scenes, especially in mixed pixels and complex backgrounds, where the noise in the segmentation results increases, making it difficult to accurately distinguish subtle changes.

Method used

By acquiring remote sensing images of the target area and multiple rounds of question-and-answer data, and using pre-trained spatial perception change detection and text description models, we fuse multiple segmentation maps, text descriptions, similar features, and difference features, adjust model parameters, and identify surface changes.

Benefits of technology

It has achieved accurate identification of changes in surface features under high-resolution remote sensing images, improving the accuracy and precision of surface change predictions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689768A_ABST
    Figure CN120689768A_ABST
Patent Text Reader

Abstract

The invention relates to the field of remote sensing data analysis, and particularly provides a method, device and equipment for identifying surface feature change and a storage medium, and the method comprises the steps: obtaining a remote sensing image of a target region and multi-round question and answer data of surface change in the target region, the multi-round question and answer data is generated according to historical actual earth surface change data in the target area; and identifying the remote sensing image and the multi-round question and answer data through a preset spatial perception change detection and text description model to obtain a change detection graph of the target area and corresponding text description. Through the method, the effect of accurately identifying the surface feature change can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of remote sensing data analysis, and in particular, to a method, apparatus, device and storage medium for identifying changes in surface features. Background Art

[0002] Change detection (CD) is a core task in remote sensing image analysis. It aims to identify surface changes, such as new buildings and road expansions, by analyzing images from different time phases. Traditional methods rely primarily on pixel-level or feature-level image segmentation techniques, combined with optical remote sensing data (such as Landsat and Sentinel-2) for analysis. These methods can effectively capture large-scale surface changes by comparing the radiometric values ​​or texture features of images. High-resolution remote sensing data provides shorter revisit periods and higher spatial resolution, which has promoted research on fine-scale change detection. However, existing remote sensing datasets typically only provide static descriptive sentences and lack multi-round question-answering data that supports contextual coherence.

[0003] However, traditional methods face multiple challenges in high-resolution imagery. First, mixed pixels and complex backgrounds increase the noise in the segmentation results, making it difficult to accurately distinguish subtle changes.

[0004] Therefore, how to accurately identify changes in surface features is a technical problem that needs to be solved. Summary of the Invention

[0005] The purpose of the embodiments of the present application is to provide a method for identifying changes in surface features. Through the technical solutions of the embodiments of the present application, the effect of accurately identifying changes in surface features can be achieved.

[0006] In a first aspect, an embodiment of the present application provides a method for identifying changes in surface features, including obtaining a remote sensing image of a target area and multi-round question-and-answer data of surface changes in the target area, wherein the multi-round question-and-answer data is generated based on historical actual surface change data in the target area; identifying the remote sensing image and the multi-round question-and-answer data through a preset spatially-aware change detection and text description model to obtain a change detection map and a corresponding text description of the target area, wherein the spatially-aware change detection and text description model is obtained by adjusting the parameters of a pre-trained language model using multiple change detection maps and multiple description features obtained by fusing multiple segmentation maps, multiple text descriptions, multiple similar features, and multiple difference features, the multiple segmentation maps and multiple text descriptions are obtained by identifying multiple data pairs using a pre-trained language model, the multiple similar features and multiple difference features are obtained by respectively comparing the multiple segmentation maps and the multiple text descriptions, and the multiple data pairs are composed of each remote sensing image in a time series remote sensing dataset, the label corresponding to each remote sensing image, and the multi-round question-and-answer data corresponding to each remote sensing image.

[0007] In the above embodiments of the present application, remote sensing images and multiple rounds of question and answer are recognized through pre-trained spatial perception change detection and text description models. By combining different data recognition results, more accurate surface change prediction results can be obtained, thereby achieving the effect of accurately identifying changes in surface features.

[0008] In some embodiments, before obtaining remote sensing images of the target area and multiple rounds of question-and-answer data on surface changes in the target area, it also includes: forming multiple data pairs from each remote sensing image in the time series remote sensing data set, the label corresponding to each remote sensing image, and the multiple rounds of question-and-answer data corresponding to each remote sensing image; identifying the multiple data pairs through a pre-trained language model to obtain multiple segmentation maps and multiple text descriptions; comparing the multiple segmentation maps and the multiple text descriptions respectively to obtain multiple similar features and multiple difference features; fusing the multiple segmentation maps, the multiple text descriptions, the multiple similar features, and the multiple difference features to obtain multiple change detection maps and multiple description features; comparing the multiple change detection maps and the multiple description features with the label corresponding to each remote sensing image, and adjusting the parameters of the pre-trained language model according to the comparison results to obtain a spatial perception change detection and text description model.

[0009] In the above embodiments of the present application, by identifying and fusing the features of remote sensing images and multi-round question-and-answer data, the parameters of the basic model pre-trained language model can be adjusted according to the final results to obtain a spatial perception change detection and text description model.

[0010] In some embodiments, multiple change detection maps and multiple descriptive features are compared with the label corresponding to each remote sensing image, and the parameters of the pre-trained language model are adjusted according to the comparison results to obtain a spatially aware change detection and text description model, including: combining multiple change detection maps, multiple descriptive features and the contextual content of multiple rounds of question and answer data corresponding to each remote sensing image to predict and output the probabilities of multiple change detection maps and multiple descriptive features; comparing the probabilities of multiple change detection maps and multiple descriptive features with the label corresponding to each remote sensing image to obtain a comparison result; according to the comparison result, adjusting the parameters of the pre-trained language model to obtain a spatially aware change detection and text description model.

[0011] In the above embodiment of the present application, the probabilities of multiple change detection maps and multiple descriptive features are compared with the labels corresponding to each remote sensing image, and the parameters of the pre-trained language model can be accurately adjusted according to the comparison results to obtain a spatially aware change detection and text description model.

[0012] In some embodiments, before forming multiple data pairs of each remote sensing image in a time series remote sensing dataset, the label corresponding to each remote sensing image, and the multiple rounds of question and answer data corresponding to each remote sensing image, it also includes: intercepting the high-resolution remote sensing image through a sliding window to obtain a time series remote sensing dataset; annotating the surface changes in each remote sensing image in the time series remote sensing dataset to obtain the label corresponding to each remote sensing image; identifying the surface change data in the time series remote sensing dataset, and answering the questions provided by the user through the surface change data, and then generating multiple rounds of question and answer data corresponding to each remote sensing image in the CoT format.

[0013] In the above embodiment of the present application, after answering the questions provided by the user through surface change data, multiple rounds of question and answer data corresponding to each remote sensing image are generated in CoT format to facilitate the subsequent model to accurately identify the question and answer data.

[0014] In some embodiments, remote sensing images and multi-round question-and-answer data are identified through a preset spatially-aware change detection and text description model to obtain a change detection map of the target area and a corresponding text description, including: identifying surface changes in remote sensing images and multi-round question-and-answer data through a spatially-aware change detection and text description model to obtain a change detection map of the target area and a corresponding text description.

[0015] In the above embodiments of the present application, the surface changes in remote sensing images and multi-round question-and-answer data are identified through spatially perceived change detection and text description models, so that the change detection map and corresponding text description of the target area can be accurately obtained.

[0016] In some embodiments, obtaining remote sensing images of a target area and multi-round question-and-answer data on surface changes within the target area includes: intercepting the remote sensing image of the target area through a sliding window to obtain the remote sensing image of the target area; inputting user questions to obtain multi-round question-and-answer data.

[0017] In the above embodiment of the present application, by intercepting the remote sensing image of the target area through a sliding window, the answer corresponding to the user's question can be identified to obtain accurate multi-round question and answer data.

[0018] In some embodiments, before obtaining remote sensing images of the target area and multiple rounds of question-and-answer data on surface changes in the target area, it also includes: obtaining dual-phase images of the target area; identifying change images in the dual-phase images to obtain remote sensing images of the target area. In the above-mentioned embodiment of the present application, remote sensing images can be accurately acquired from the changing images in the dual-temporal images.

[0019] In a second aspect, an embodiment of the present application provides a device for identifying changes in surface features, comprising: An acquisition module is used to obtain remote sensing images of the target area and multi-round question-and-answer data on surface changes in the target area, wherein the multi-round question-and-answer data is generated based on historical actual surface change data in the target area; The recognition module is used to identify remote sensing images and multi-round question and answer data through a preset spatially aware change detection and text description model to obtain a change detection map and a corresponding text description of the target area. The spatially aware change detection and text description model is obtained by adjusting the parameters of a pre-trained language model by fusing multiple segmentation maps, multiple text descriptions, multiple similar features, and multiple difference features to obtain multiple change detection maps and multiple description features. The multiple segmentation maps and multiple text descriptions are obtained by identifying multiple data pairs using a pre-trained language model. The multiple similar features and multiple difference features are obtained by respectively comparing the multiple segmentation maps and multiple text descriptions. The multiple data pairs are composed of each remote sensing image in a time series remote sensing dataset, the label corresponding to each remote sensing image, and the multi-round question and answer data corresponding to each remote sensing image.

[0020] Optionally, the device further includes: a training module configured to, before the acquisition module acquires the remote sensing images of the target area and the multi-round question-and-answer data of the surface changes in the target area, form a plurality of data pairs from each remote sensing image in the time series remote sensing dataset, the label corresponding to each remote sensing image, and the multi-round question-and-answer data corresponding to each remote sensing image; Identify multiple data pairs through a pre-trained language model to obtain multiple segmentation maps and multiple text descriptions; Compare multiple segmentation images and multiple text descriptions to obtain multiple similar features and multiple different features; Fusion of multiple segmentation maps, multiple text descriptions, multiple similarity features, and multiple difference features to obtain multiple change detection maps and multiple description features; Multiple change detection maps and multiple description features are compared with the labels corresponding to each remote sensing image. The parameters of the pre-trained language model are adjusted according to the comparison results to obtain a spatially aware change detection and text description model.

[0021] Optionally, the acquisition module is specifically configured to: Combining multiple change detection maps, multiple descriptive features, and the context of multiple rounds of question-answering data corresponding to each remote sensing image, the probabilities of multiple change detection maps and multiple descriptive features are predicted and output; Compare the multiple change detection maps and the probabilities of the multiple description features with the labels corresponding to each remote sensing image to obtain comparison results; According to the comparison results, the parameters of the pre-trained language model are adjusted to obtain a spatially aware change detection and text description model.

[0022] Optionally, the device further includes: a generation module configured to intercept the high-resolution remote sensing images through a sliding window before the training module forms multiple data pairs from each remote sensing image, the label corresponding to each remote sensing image, and the multiple rounds of question-answering data corresponding to each remote sensing image in the time series remote sensing dataset to obtain a time series remote sensing dataset; By annotating the surface changes in each remote sensing image in the time series remote sensing dataset, the label corresponding to each remote sensing image is obtained; Identify surface change data in time-series remote sensing datasets, answer user questions based on the surface change data, and generate multiple rounds of question-answering data corresponding to each remote sensing image in the CoT format.

[0023] Optionally, the prediction module is specifically used to: Through the spatial perception change detection and text description model, the surface changes in remote sensing images and multi-round question-answering data are identified, and the change detection map and corresponding text description of the target area are obtained.

[0024] Optionally, the acquisition module is specifically used to: The remote sensing image of the target area is intercepted by a sliding window to obtain the remote sensing image of the target area; Enter the user's question and obtain multiple rounds of question-answering data.

[0025] Optionally, the device further includes: an identification module configured to obtain a dual-temporal image of the target area before the acquisition module acquires the remote sensing image of the target area and the multi-round question-and-answer data of the surface changes in the target area; Identify the changing images in the dual-temporal images and obtain the remote sensing images of the target area.

[0026] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor and a memory, wherein the memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, the steps in the method provided in the first aspect above are executed.

[0027] In a fourth aspect, an embodiment of the present application provides a readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps in the method provided in the first aspect are executed.

[0028] Other features and advantages of the present application will be described in the subsequent description, and in part will become apparent from the description, or will be understood by practicing the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.

[0030] Figure 1 A flow chart of a method for identifying changes in surface features provided in an embodiment of the present application; Figure 2 A flowchart of an implementation method for identifying changes in surface features provided in an embodiment of the present application; Figure 3 A schematic block diagram of a device for identifying changes in surface features provided in an embodiment of the present application; Figure 4 A schematic block diagram of the structure of a device for identifying changes in surface features provided in an embodiment of the present application. DETAILED DESCRIPTION

[0031] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work fall within the scope of protection of the present application.

[0032] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and should not be understood as indicating or implying relative importance.

[0033] First, some of the terms involved in the embodiments of the present application are explained to facilitate understanding by those skilled in the art.

[0034] The Chain-of-Thought (CoT) format is a technique for improving the reasoning capabilities of large language models (LLMs) and explaining their thought processes.

[0035] Dual-temporal imaging refers to two or more images of the same scene taken at different points in time. These images are usually taken at different times to capture information about changes over time.

[0036] This application is applied to the scenario of remote sensing data and text recognition. The specific scenario is to generate CoT format question-answering data by creating and designing a dual-temporal remote sensing dataset, combining pre-trained models and fine-tuning strategies to achieve accurate recognition of spatial changes and context-coherent semantic text description, supporting large-scale surface prediction.

[0037] Change detection (CD) is a core task in remote sensing image analysis. It aims to identify surface changes, such as new buildings and road expansions, by analyzing images from different time phases. Traditional methods rely primarily on pixel-level or feature-level image segmentation techniques, combined with optical remote sensing data (such as Landsat and Sentinel-2). These methods effectively capture large-scale surface changes by comparing image radiometric values ​​or texture features. High-resolution remote sensing data offers shorter revisit periods and higher spatial resolution, driving research in fine-scale change detection. However, existing remote sensing datasets typically provide only static descriptive sentences and lack contextually coherent multi-round question-and-answer data. However, traditional methods face multiple challenges in high-resolution imagery. First, mixed pixels and complex backgrounds increase the noise in the segmentation results, making it difficult to accurately distinguish subtle changes.

[0038] To this end, the present application obtains remote sensing images of a target area and multi-round question-and-answer data on surface changes in the target area, wherein the multi-round question-and-answer data is generated based on historical actual surface change data in the target area; the remote sensing images and multi-round question-and-answer data are identified by a preset spatially aware change detection and text description model to obtain a change detection map and a corresponding text description of the target area, wherein the spatially aware change detection and text description model is obtained by adjusting the parameters of a pre-trained language model by fusing multiple segmentation maps, multiple text descriptions, multiple similar features, and multiple difference features to obtain multiple change detection maps and multiple description features; the multiple segmentation maps and multiple text descriptions are obtained by identifying multiple data pairs using a pre-trained language model; the multiple similar features and multiple difference features are obtained by comparing multiple segmentation maps and multiple text descriptions respectively; the multiple data pairs are composed of each remote sensing image in a time series remote sensing dataset, the label corresponding to each remote sensing image, and the multi-round question-and-answer data corresponding to each remote sensing image. By identifying remote sensing images and multi-round question-and-answer data using a pre-trained spatially aware change detection and text description model, a more accurate surface change prediction result can be obtained by combining the recognition results of different data, thereby achieving the effect of accurately identifying changes in surface features.

[0039] In the embodiment of the present application, the execution entity may be a surface feature change identification device in a surface feature change identification system. In actual applications, the surface feature change identification device may be an electronic device such as a terminal device and a server, and no limitation is made here.

[0040] The following combination Figure 1 The method for identifying changes in surface features in an embodiment of the present application is described in detail.

[0041] Please see Figure 1 , Figure 1 A flow chart of a method for identifying changes in surface features provided in an embodiment of the present application is shown in FIG. Figure 1 The methods shown for identifying changes in surface characteristics include: Step 110: Acquire remote sensing images of the target area and multi-round question-and-answer data on surface changes in the target area.

[0042] The multi-round question-and-answer data is generated based on historical data on actual surface changes within the target area. It can also be acquired through multi-round conversations between the robot's question-and-answer model and the user. The target area can be any surface area, including mountains, farms, urban buildings, and rural roads.

[0043] In some embodiments of the present application, before acquiring remote sensing images of a target area and multiple rounds of question-and-answer data on surface changes within the target area, Figure 1The method shown also includes: forming multiple data pairs from each remote sensing image in a time series remote sensing dataset, the label corresponding to each remote sensing image, and multiple rounds of question-and-answer data corresponding to each remote sensing image; identifying the multiple data pairs through a pre-trained language model to obtain multiple segmentation maps and multiple text descriptions; comparing the multiple segmentation maps and the multiple text descriptions respectively to obtain multiple similar features and multiple difference features; fusing the multiple segmentation maps, the multiple text descriptions, the multiple similar features, and the multiple difference features to obtain multiple change detection maps and multiple description features; comparing the multiple change detection maps and the multiple description features with the label corresponding to each remote sensing image, and adjusting the parameters of the pre-trained language model according to the comparison results to obtain a spatial perception change detection and text description model.

[0044] In this application, by identifying and fusing features from remote sensing images and multi-round question-and-answer data, the parameters of the pre-trained language model in the base model are adjusted based on the final results to produce a spatially aware change detection and text description model. The comparison results include differences and loss calculations.

[0045] The time-series remote sensing dataset includes multiple remote sensing images of the target area captured by different satellites. The labels associated with each remote sensing image are obtained based on annotations by relevant personnel. Each data set includes the remote sensing image, the corresponding label, and multiple rounds of question-and-answer data. Similar features are obtained through similarity calculation, and the remaining features are difference features.

[0046] Optionally, before forming multiple data pairs including each remote sensing image in the time series remote sensing dataset, the label corresponding to each remote sensing image, and the multiple rounds of question and answer data corresponding to each remote sensing image, the high-resolution time series remote sensing dataset is also preprocessed to obtain each remote sensing image, the label corresponding to each remote sensing image, and the multiple rounds of question and answer data corresponding to each remote sensing image.

[0047] Specifically: First, this method is based on an original high-resolution time-series remote sensing dataset, which includes dual-phase images (A / B phases), image labels (labeling the change area and category), and CoT format question-answering data. The dual-phase images are normalized using the following calculation formula: ; in, is the original image value of channel c at position (h, w), μc and σc are the mean and standard deviation of channel c, respectively, and I′(c, h, w) is the normalized image value. Subsequently, the segmentation label pixel values ​​are mapped to categories, and the categories are divided into buildings, roads, and unchanged parts. The calculation formula is as follows: ; in, is the original pixel value, L(h,w) is the class label, and the class label values ​​2, 1, and 0 represent buildings, roads, and unchanged parts respectively. Optionally, the label corresponding to each remote sensing image is obtained in the following way: the sample annotation mainly consists of two parts. The first part is a mask map of the location of the changed area, which is manually and visually annotated; the second part is a text CoT description annotation. First, create CoT format question and answer data for the dual-temporal image, including multiple rounds of question and answer pairs (such as "Have these two images changed?" "Yes." "What is the type of the newly added changed area?" "A new building has been added." "Where is the building?" "In the middle of the image.") to ensure that the questions and answers are contextually coherent. Then use the pre-trained language model BERT to encode the CoT question and answer data and generate a token word sequence. The calculation formula is as follows: ; Among them, c j is the jth answer, is a word sequence and , The effective length is the maximum sequence length. The role of the effective length is to help the model focus only on meaningful tokens when processing sequence data and ignore the padding, thereby improving the efficiency and accuracy of model training and prediction. The effective length calculation formula is as follows: ; in, is the effective length of the j-th answer, and pad is the padding marker. The CoT question and answer are then paired with the bi-temporal image to ensure that the semantics of the question and the changed area are consistent. Finally, the dataset is divided into training set, validation set, and test set for subsequent model training, optimization, and testing.

[0048] Optionally, multiple data pairs are identified using a pre-trained language model to obtain multiple segmentation maps and multiple text descriptions; the multiple segmentation maps and multiple text descriptions are compared to obtain multiple similar features and multiple different features, including: Identify the changed areas of the dual-temporal image and generate a segmentation map; the change description branch (Change Caption, CC) focuses on generating semantic text descriptions related to the changed parts and supports CoT question answering. As input, a pre-trained visual model is used to extract multi-scale features. The calculation formula is as follows: ; in, Phase images The multi-scale features extracted and , N is the number of feature layers.

[0049] Optionally, multiple segmentation maps, multiple text descriptions, multiple similar features and multiple difference features are fused to obtain multiple change detection maps and multiple description features, including: The CoT problem P = {p1, p2, ...} and the answer C = {c1, c2, ...} are input. In the change detection branch, feature differences and similarities are calculated. Feature differences highlight the pixel-level changes between the two temporal images, and cosine similarity measures the similarity between feature vectors. The combination of the two provides comprehensive information for subsequent Transformer fusion, helping the model accurately locate the changed area. The calculation formula is as follows: ; Among them, Di is the difference feature, Si is the cosine similarity, and Transformer is used to fuse the features and project them into the segmentation space. The calculation formula is as follows: : in, is the fusion feature of the i-th layer. Transformer models Di and Si through a multi-head self-attention mechanism to capture the changing spatiotemporal dependencies and then transforms the multi-scale features into Spliced ​​and fused through convolution operation , thereby generating a predicted segmentation map, the calculation formula is as follows: : in, To predict the segmentation map, K is the number of categories, Conv uses a 1×1 convolution kernel to Mapped to the category space, Upsample restores the feature map to the original resolution through bilinear interpolation.

[0050] Encode the CoT question and answer in the change description branch, and the calculation formula is as follows: : The encoding result contains the semantic information of the text and is used for cross-modal feature projection: (10)Equation Chapter (Next) Section 1 (11)Equation Chapter (Next) Section 1 : in, , , It is the projection layer, which outputs the description features , and retain context information, output segmentation map and descriptive features .

[0051] Optionally, multiple change detection maps and multiple description features are compared with the labels corresponding to each remote sensing image, and parameters of a pre-trained language model are adjusted based on the comparison results to obtain a spatially aware change detection and text description model including: In the sequence generation stage, the features are described and historical issues As input, use the Transformer decoder to generate a sequence, combined with the historical context, the calculation formula is as follows: ; in, , V is the vocabulary size, and text is generated through beam search. The calculation formula is as follows: ; Output predicted probability (training) or text sequence (Inference),In the training phase, the loss function is defined, and the change detection loss is as follows: ; Among them, C is the number of categories, H, W are the segmentation map sizes, is the predicted probability, is the true label.

[0052] The change description loss is as follows: ; Where L is the sequence length, V is the vocabulary size, is the real token at the lth position, is the predicted probability.

[0053] The joint loss formula is shown in 17: Equation Chapter (Next) Section 1 ; Using Adam optimizer, learning rate , the gradient accumulation step is as follows: ; Among them, B is the batch size, initial training, joint training In the future, we plan to expand the dataset to hundreds of images, add different forms of CoT question-and-answer data, and optimize model performance.

[0054] In some embodiments of the present application, multiple change detection maps and multiple descriptive features are compared with the label corresponding to each remote sensing image, and the parameters of the pre-trained language model are adjusted according to the comparison results to obtain a spatially aware change detection and text description model, including: combining multiple change detection maps, multiple descriptive features and the contextual content of multiple rounds of question and answer data corresponding to each remote sensing image, predicting and outputting the probabilities of multiple change detection maps and multiple descriptive features; comparing the probabilities of multiple change detection maps and multiple descriptive features with the label corresponding to each remote sensing image, respectively, to obtain a comparison result; and adjusting the parameters of the pre-trained language model according to the comparison result to obtain a spatially aware change detection and text description model.

[0055] In the above process, the present application compares the probabilities of multiple change detection maps and multiple descriptive features with the labels corresponding to each remote sensing image, and can accurately adjust the parameters of the pre-trained language model according to the comparison results to obtain a spatially aware change detection and text description model.

[0056] In some embodiments of the present application, before forming multiple data pairs from each remote sensing image in the time series remote sensing dataset, the label corresponding to each remote sensing image, and the multiple rounds of question-answering data corresponding to each remote sensing image, Figure 1 The method shown also includes: intercepting high-resolution remote sensing images through a sliding window to obtain a time series remote sensing dataset; annotating the surface changes in each remote sensing image in the time series remote sensing dataset to obtain a label corresponding to each remote sensing image; identifying surface change data in the time series remote sensing dataset, and answering questions provided by the user based on the surface change data, and then generating multiple rounds of question and answer data corresponding to each remote sensing image in the CoT format.

[0057] In the above process, this application uses surface change data to answer questions provided by users and then generates multiple rounds of question and answer data corresponding to each remote sensing image in CoT format to facilitate the subsequent model to accurately identify the question and answer data.

[0058] The sliding window can be set as required, and the change data can be obtained by comparing the current surface image with the historical surface image.

[0059] In some embodiments of the present application, remote sensing images of a target area and multi-round question-and-answer data of surface changes within the target area are obtained, including: intercepting the remote sensing image of the target area through a sliding window to obtain the remote sensing image of the target area; inputting the user's questions to obtain multi-round question-and-answer data.

[0060] In the above process, the present application intercepts the remote sensing image of the target area through a sliding window to identify the answer corresponding to the user's question and obtain accurate multi-round question and answer data.

[0061] In some embodiments of the present application, before acquiring remote sensing images of a target area and multiple rounds of question-and-answer data on surface changes within the target area, Figure 1 The method shown also includes: acquiring dual-temporal images of the target area; identifying the change images in the dual-temporal images to obtain a remote sensing image of the target area. In the above process, the present application can accurately obtain remote sensing images from the change images in the dual-phase images.

[0062] Among them, when identifying the above-mentioned images, videos and changed images, the data information in the image can be obtained after identification based on the image recognition model obtained by training the basic model according to the existing model or historical image data.

[0063] Step 120: The remote sensing image and the multi-round question-answering data are identified by using a preset spatially aware change detection and text description model to obtain a change detection map of the target area and a corresponding text description.

[0064] Among them, the spatial perception change detection and text description model is obtained by adjusting the parameters of the pre-trained language model by fusing multiple change detection maps and multiple description features obtained by fusing multiple segmentation maps, multiple text descriptions, multiple similar features and multiple difference features. Multiple segmentation maps and multiple text descriptions are obtained by identifying multiple data pairs through the pre-trained language model. Multiple similar features and multiple difference features are obtained by respectively comparing multiple segmentation maps and multiple text descriptions. Multiple data pairs are composed of each remote sensing image in the time series remote sensing dataset, the label corresponding to each remote sensing image and multiple rounds of question and answer data corresponding to each remote sensing image.

[0065] In some embodiments of the present application, remote sensing images and multi-round question-and-answer data are identified through a preset spatially-aware change detection and text description model to obtain a change detection map of the target area and a corresponding text description, including: identifying surface changes in remote sensing images and multi-round question-and-answer data through a spatially-aware change detection and text description model to obtain a change detection map of the target area and a corresponding text description.

[0066] In the above process, this application uses spatially perceived change detection and text description models to identify surface changes in remote sensing images and multi-round question-and-answer data, and can accurately obtain change detection maps and corresponding text descriptions of the target area.

[0067] In the above Figure 1In the process shown, the present application obtains remote sensing images of the target area and multi-round question-and-answer data of surface changes in the target area, wherein the multi-round question-and-answer data is generated based on historical actual surface change data in the target area; the remote sensing images and multi-round question-and-answer data are identified by a preset spatially aware change detection and text description model to obtain a change detection map and a corresponding text description of the target area, wherein the spatially aware change detection and text description model is obtained by adjusting the parameters of a pre-trained language model by fusing multiple segmentation maps, multiple text descriptions, multiple similar features, and multiple difference features to obtain multiple change detection maps and multiple description features, the multiple segmentation maps and multiple text descriptions are obtained by identifying multiple data pairs using a pre-trained language model, the multiple similar features and multiple difference features are obtained by respectively comparing multiple segmentation maps and multiple text descriptions, and the multiple data pairs are composed of each remote sensing image in the time series remote sensing dataset, the label corresponding to each remote sensing image, and the multi-round question-and-answer data corresponding to each remote sensing image. By pre-training spatially aware change detection and text description models to identify remote sensing images and multiple rounds of question and answer, we can obtain more accurate surface change prediction results by combining the recognition results of different data, thereby achieving the effect of accurately identifying changes in surface features.

[0068] The following combination Figure 2 The implementation method of identifying changes in surface features in an embodiment of the present application is described in detail.

[0069] Please see Figure 2 , Figure 2 A flowchart of an implementation method for identifying changes in surface features provided in an embodiment of the present application is shown in FIG. Figure 2 The implementation method shown for identifying changes in surface characteristics includes: Step 210: Image recognition.

[0070] Specifically: Identify the remote sensing image of the target area and obtain image features.

[0071] Step 220: Question and answer recognition.

[0072] Specifically: Identify multiple rounds of question-answering data corresponding to each remote sensing image and obtain text features.

[0073] Step 230: Feature fusion.

[0074] Specifically: image features and text features as well as similar features and difference features corresponding to the segmentation map and text description are integrated to obtain multiple change detection maps and multiple description features.

[0075] Step 240: Surface identification.

[0076] Specifically: multiple change detection images and multiple description features are converted into prediction images and text descriptions to obtain recognition results.

[0077] also, Figure 2 The specific methods and steps shown can be found in Figure 1 The method shown here will not be described in detail.

[0078] Previous article passed Figure 1-Figure 2 The method of identifying changes in surface features is described below. Figure 3-Figure 4 Describe a device for identifying changes in surface features.

[0079] Please refer to Figure 3 , is a schematic block diagram of a device 300 for identifying changes in surface features provided in an embodiment of the present application. The device 300 may be a module, program segment or code on an electronic device. The device 300 is similar to the above-mentioned Figure 1 The method embodiment corresponds to the embodiment that can be executed Figure 1 The various steps involved in the method embodiment and the specific functions of the device 300 can be found in the description below. To avoid repetition, detailed description is appropriately omitted here.

[0080] Optionally, the device 300 includes: An acquisition module 310 is configured to acquire remote sensing images of a target area and multi-round question-and-answer data on surface changes within the target area, wherein the multi-round question-and-answer data is generated based on historical actual surface change data within the target area; The recognition module 320 is used to recognize remote sensing images and multi-round question and answer data through a preset spatially aware change detection and text description model to obtain a change detection map and a corresponding text description of the target area, wherein the spatially aware change detection and text description model is obtained by adjusting the parameters of a pre-trained language model by fusing multiple segmentation maps, multiple text descriptions, multiple similar features, and multiple difference features to obtain multiple change detection maps and multiple description features; the multiple segmentation maps and multiple text descriptions are obtained by recognizing multiple data pairs using a pre-trained language model; the multiple similar features and multiple difference features are obtained by respectively comparing the multiple segmentation maps and the multiple text descriptions; and the multiple data pairs are composed of each remote sensing image in a time series remote sensing dataset, the label corresponding to each remote sensing image, and the multi-round question and answer data corresponding to each remote sensing image.

[0081] Optionally, the device further includes: A training module is used for the acquisition module to form multiple data pairs from each remote sensing image in the time series remote sensing data set, the label corresponding to each remote sensing image, and the multiple rounds of question and answer data corresponding to each remote sensing image before acquiring the remote sensing image of the target area and the multiple rounds of question and answer data of the surface changes in the target area; recognize the multiple data pairs through a pre-trained language model to obtain multiple segmentation maps and multiple text descriptions; compare the multiple segmentation maps and the multiple text descriptions respectively to obtain multiple similar features and multiple difference features; fuse the multiple segmentation maps, the multiple text descriptions, the multiple similar features, and the multiple difference features to obtain multiple change detection maps and multiple description features; compare the multiple change detection maps and the multiple description features with the label corresponding to each remote sensing image, adjust the parameters of the pre-trained language model according to the comparison results, and obtain a spatial perception change detection and text description model.

[0082] Optionally, the acquisition module is specifically configured to: Combining multiple change detection maps, multiple descriptive features, and the contextual content of multiple rounds of question-and-answer data corresponding to each remote sensing image, the probabilities of multiple change detection maps and multiple descriptive features are predicted and output; the probabilities of multiple change detection maps and multiple descriptive features are compared with the labels corresponding to each remote sensing image to obtain comparison results; based on the comparison results, the parameters of the pre-trained language model are adjusted to obtain a spatially aware change detection and text description model.

[0083] Optionally, the device further includes: A generation module is used for the training module to intercept high-resolution remote sensing images through a sliding window to obtain a time-series remote sensing dataset before forming multiple data pairs from each remote sensing image in the time-series remote sensing dataset, the label corresponding to each remote sensing image, and the multi-round question-and-answer data corresponding to each remote sensing image; to obtain the label corresponding to each remote sensing image by marking the surface changes in each remote sensing image in the time-series remote sensing dataset; to identify the surface change data in the time-series remote sensing dataset, and to generate multi-round question-and-answer data corresponding to each remote sensing image in the CoT format after answering the questions provided by the user using the surface change data.

[0084] Optionally, the prediction module is specifically used to: Through the spatial perception change detection and text description model, the surface changes in remote sensing images and multi-round question-answering data are identified, and the change detection map and corresponding text description of the target area are obtained.

[0085] Optionally, the acquisition module is specifically used to: The remote sensing image of the target area is captured by sliding the window to obtain the remote sensing image of the target area; the user's questions are input to obtain multiple rounds of question and answer data.

[0086] Optionally, the device further includes: The recognition module is used to obtain a dual-phase image of the target area before the acquisition module obtains the remote sensing image of the target area and multiple rounds of question-and-answer data on surface changes in the target area; identify the change image in the dual-phase image to obtain the remote sensing image of the target area.

[0087] Please refer to Figure 4 This is a schematic block diagram of the structure of a device for identifying surface feature changes provided in an embodiment of the present application. The device may include a memory 410 and a processor 420. Optionally, the device may also include: a communication interface 430 and a communication bus 440. The device is similar to the above Figure 1 The method embodiment corresponds to the embodiment that can be executed Figure 1 The various steps involved in the method embodiment and the specific functions of the device can be found in the description below.

[0088] Specifically, the memory 410 is used to store computer-readable instructions.

[0089] Processor 420 is used to process the readable instructions stored in the memory and execute Figure 1 The steps in the method.

[0090] The communication interface 430 is used for signaling or data communication with other node devices, for example, for communication with a server or terminal, or for communication with other device nodes, but the embodiments of the present application are not limited thereto.

[0091] The communication bus 440 is used to realize direct connection and communication among the above components.

[0092] Among them, the communication interface 430 of the device in the embodiment of the present application is used to communicate signaling or data with other node devices. The memory 410 can be a high-speed RAM memory or a non-volatile memory (non-volatile memory), such as at least one disk memory. The memory 410 can also be at least one storage device located away from the aforementioned processor. The memory 410 stores computer-readable instructions. When the computer-readable instructions are executed by the processor 420, the electronic device executes the above-mentioned Figure 1The method process shown. The processor 420 can be used on the device 300 and is used to perform the functions of the present application. For example, the above-mentioned processor 420 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components, but the embodiments of the present application are not limited thereto.

[0093] The embodiment of the present application further provides a readable storage medium, wherein when the computer program is executed by a processor, Figure 1 The method process in the illustrated method embodiment is performed by the electronic device.

[0094] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working process of the device described above can refer to the corresponding process in the aforementioned method, and will not be described in detail here.

[0095] In summary, the embodiments of the present application provide a method, apparatus, device and storage medium for identifying changes in surface features. The method includes obtaining a remote sensing image of a target area and multi-round question-and-answer data of surface changes in the target area, wherein the multi-round question-and-answer data is generated based on historical actual surface change data in the target area; identifying the remote sensing image and the multi-round question-and-answer data through a preset spatially aware change detection and text description model to obtain a change detection map and a corresponding text description of the target area, wherein the spatially aware change detection and text description model is obtained by adjusting the parameters of a pre-trained language model by fusing multiple segmentation maps, multiple text descriptions, multiple similar features and multiple difference features to obtain multiple change detection maps and multiple description features, multiple segmentation maps and multiple text descriptions are obtained by identifying multiple data pairs through a pre-trained language model, multiple similar features and multiple difference features are obtained by respectively comparing multiple segmentation maps and multiple text descriptions, and multiple data pairs are composed of each remote sensing image in a time series remote sensing dataset, a label corresponding to each remote sensing image and multi-round question-and-answer data corresponding to each remote sensing image. This method can achieve the effect of accurately identifying changes in surface features.

[0096] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to the multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or can be implemented using a combination of dedicated hardware and computer instructions.

[0097] In addition, the functional modules in each embodiment of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0098] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard drives, read-only memories (ROM), random access memories (RAM), magnetic disks or optical disks.

[0099] The foregoing is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included within the scope of protection of the present application. It should be noted that similar reference numerals and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined or explained in subsequent figures.

[0100] The above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.

[0101] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

Claims

1. A method for identifying changes in surface features, characterized in that: include: Acquire a remote sensing image of a target area and multi-round question-and-answer data on surface changes within the target area, wherein the multi-round question-and-answer data is generated based on historical actual surface change data within the target area; The remote sensing image and the multiple rounds of question-and-answer data are identified by a preset spatially-aware change detection and text description model to obtain a change detection map and a corresponding text description of the target area, wherein the spatially-aware change detection and text description model is obtained by adjusting the parameters of a pre-trained language model using multiple change detection maps and multiple description features obtained by fusing multiple segmentation maps, multiple text descriptions, multiple similar features, and multiple difference features. The multiple segmentation maps and the multiple text descriptions are obtained by identifying multiple data pairs using the pre-trained language model, the multiple similar features and the multiple difference features are obtained by respectively comparing the multiple segmentation maps and the multiple text descriptions, and the multiple data pairs are composed of each remote sensing image in a time series remote sensing dataset, the label corresponding to each remote sensing image, and the multiple rounds of question-and-answer data corresponding to each remote sensing image.

2. The method according to claim 1, characterized in that Before acquiring the remote sensing image of the target area and the multi-round question-and-answer data of the surface changes in the target area, the method further includes: Each remote sensing image in the time series remote sensing dataset, a label corresponding to each remote sensing image, and multiple rounds of question-answering data corresponding to each remote sensing image form the multiple data pairs; Identify the multiple data pairs using the pre-trained language model to obtain the multiple segmentation maps and the multiple text descriptions; Comparing the plurality of segmentation images and the plurality of text descriptions respectively to obtain the plurality of similar features and the plurality of different features; fusing the plurality of segmentation maps, the plurality of text descriptions, the plurality of similar features, and the plurality of difference features to obtain the plurality of change detection maps and the plurality of description features; The multiple change detection maps and the multiple description features are compared with the labels corresponding to each remote sensing image, and the parameters of the pre-trained language model are adjusted according to the comparison results to obtain the spatially aware change detection and text description model.

3. The method according to claim 2, characterized in that The step of comparing the multiple change detection maps and the multiple description features with the labels corresponding to each remote sensing image, and adjusting the parameters of the pre-trained language model according to the comparison results to obtain the spatially aware change detection and text description model includes: combining the multiple change detection maps, the multiple descriptive features, and contextual content of multiple rounds of question-answering data corresponding to each remote sensing image to predict and output probabilities of the multiple change detection maps and the multiple descriptive features; Comparing the multiple change detection maps and the multiple probabilities of the descriptive features with the labels corresponding to each remote sensing image to obtain a comparison result; According to the comparison result, the parameters of the pre-trained language model are adjusted to obtain the spatial perception change detection and text description model.

4. The method according to claim 2, characterized in that Before forming the plurality of data pairs from each remote sensing image in the time series remote sensing dataset, the label corresponding to each remote sensing image, and the multiple rounds of question-answering data corresponding to each remote sensing image, the method further includes: The high-resolution remote sensing image is intercepted by a sliding window to obtain the time series remote sensing dataset; Obtaining a label corresponding to each remote sensing image by annotating the surface changes in each remote sensing image in the time series remote sensing dataset; The surface change data in the time series remote sensing data set is identified, and after the user's questions are answered using the surface change data, multiple rounds of question-answering data corresponding to each remote sensing image are generated in a CoT format.

5. The method according to any one of claims 1 to 3, characterized in that The remote sensing image and the multiple rounds of question-answering data are identified by using a preset spatially aware change detection and text description model to obtain a change detection map and a corresponding text description of the target area, including: The surface changes in the remote sensing image and the multi-round question-answering data are identified by the spatially-aware change detection and text description model to obtain a change detection map of the target area and the corresponding text description.

6. The method according to any one of claims 1 to 3, characterized in that The acquiring of remote sensing images of a target area and multiple rounds of question-and-answer data on surface changes within the target area includes: intercepting the remote sensing image of the target area through a sliding window to obtain the remote sensing image of the target area; Input the user's question and obtain the multiple rounds of question-answering data.

7. The method according to any one of claims 1 to 3, characterized in that Before acquiring the remote sensing image of the target area and the multi-round question-and-answer data of the surface changes in the target area, the method further includes: Acquiring a dual-phase image of the target area; The change image in the dual-temporal image is identified to obtain a remote sensing image of the target area.

8. A device for identifying changes in surface features, characterized in that: include: an acquisition module, configured to acquire a remote sensing image of a target area and multi-round question-and-answer data on surface changes within the target area, wherein the multi-round question-and-answer data is generated based on historical actual surface change data within the target area; An identification module is used to identify the remote sensing image and the multiple rounds of question and answer data through a preset spatially aware change detection and text description model to obtain a change detection map and a corresponding text description of the target area, wherein the spatially aware change detection and text description model is obtained by adjusting the parameters of a pre-trained language model by fusing multiple segmentation maps, multiple text descriptions, multiple similar features, and multiple difference features to obtain multiple change detection maps and multiple description features; the multiple segmentation maps and the multiple text descriptions are obtained by identifying multiple data pairs through the pre-trained language model; the multiple similar features and the multiple difference features are obtained by respectively comparing the multiple segmentation maps and the multiple text descriptions; and the multiple data pairs are composed of each remote sensing image in a time series remote sensing dataset, the label corresponding to each remote sensing image, and the multiple rounds of question and answer data corresponding to each remote sensing image.

9. An electronic device, characterized in that: include: A memory and a processor, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the steps of the method according to any one of claims 1 to 7 are executed.

10. A computer-readable storage medium, characterized in that include: A computer program, when running on a computer, causes the computer to perform the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Remote sensing cultivated land change detection method and system considering phenological characteristics

    CN111768101A

  • Remote sensing image processing method and system based on space-based remote sensing model, electronic equipment and medium

    CN119152373A

  • Building target damage assessment method based on multi-modal remote sensing data

    CN119810523A

  • Systems, methods, devices, and platforms for industrial internet of things

    WO2024155584A1