A method, apparatus, device, and storage medium for identifying changes in surface features

By combining multi-round question-and-answer data and a pre-trained language model in high-resolution remote sensing imagery, change detection maps and text descriptions are generated, solving the accuracy problem of identifying subtle surface changes in traditional methods and achieving higher-precision detection of surface feature changes.

CN120689768BActive Publication Date: 2025-12-30BEIJING NORMAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510826640.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-12-30
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

Traditional remote sensing image analysis methods struggle to accurately identify subtle changes on the land surface in high-resolution scenes, especially when segmentation results are noisy due to mixed pixels and complex backgrounds, making it difficult to accurately distinguish subtle changes.

Method used

By acquiring remote sensing images and multi-round question-and-answer data of the target area, the recognition results of multiple data are fused using a pre-trained language model. Multiple data pairs are identified using the pre-trained language model, and multiple change detection maps and text descriptions are generated by combining the multi-round question-and-answer data. The parameters of the pre-trained language model are adjusted to achieve more accurate prediction of land surface changes.

Benefits of technology

It enables more accurate identification of surface feature changes under high-resolution remote sensing imagery, improving the accuracy and reliability of surface change detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689768B_ABST
    Figure CN120689768B_ABST
Patent Text Reader

Abstract

The application relates to the field of remote sensing data analysis, and specifically provides a method and device for identifying surface feature changes, equipment and a storage medium, the method comprising the following steps: acquiring remote sensing images of a target area and multi-round question-and-answer data of surface changes in the target area, wherein the multi-round question-and-answer data is generated according to historical actual surface change data in the target area; identifying the remote sensing images and the multi-round question-and-answer data through a preset spatial perception change detection and text description model to obtain a change detection map of the target area and corresponding text description. Through the method, the effect of accurately identifying surface feature changes can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of remote sensing data analysis, and more specifically, to a method, apparatus, device, and storage medium for identifying changes in surface features. Background Technology

[0002] Change detection (CD) is a core task in remote sensing image analysis, aiming to identify surface changes, such as new building construction and road expansion, by analyzing images from different time periods. Traditional methods primarily rely on pixel-level or feature-level image segmentation techniques combined with optical remote sensing data (such as Landsat and Sentinel-2). These methods effectively capture large-scale surface changes by comparing the radiometric or textural features of the images. High-resolution remote sensing data provides shorter revisit periods and higher spatial resolution, driving research into fine-scale change detection. However, existing remote sensing datasets typically only provide static descriptive statements and lack multi-turn question-and-answer data that supports coherent context.

[0003] However, traditional methods face multiple challenges in high-resolution imagery scenarios. First, mixed pixels and complex backgrounds increase noise in the segmentation results, making it difficult to accurately distinguish subtle changes.

[0004] Therefore, accurately identifying changes in surface features is a technical problem that needs to be solved. Summary of the Invention

[0005] The purpose of this application is to provide a method for identifying changes in land surface features. The technical solution of this application can achieve the effect of accurately identifying changes in land surface features.

[0006] In a first aspect, embodiments of this application provide a method for identifying changes in land surface features, comprising: acquiring remote sensing images of a target area and multi-round question-and-answer data on land surface changes within the target area, wherein the multi-round question-and-answer data is generated based on historical actual land surface change data within the target area; identifying the remote sensing images and multi-round question-and-answer data using a preset spatially perceived change detection and text description model to obtain a change detection map and corresponding text description of the target area, wherein the spatially perceived change detection and text description model is obtained by adjusting the parameters of a pre-trained language model by fusing multiple segmentation maps, multiple text descriptions, multiple similar features, and multiple difference features; the multiple segmentation maps and multiple text descriptions are obtained by identifying multiple data pairs using the pre-trained language model; the multiple similar features and multiple difference features are obtained by comparing multiple segmentation maps and multiple text descriptions respectively; and the multiple data pairs are composed of each remote sensing image in a time-series remote sensing dataset, the label corresponding to each remote sensing image, and the multi-round question-and-answer data corresponding to each remote sensing image.

[0007] In the above embodiments, this application uses a pre-trained spatial perception change detection and text description model to identify remote sensing images and multi-turn question answering. By combining the identification results of different data, a more accurate prediction result of land surface change can be obtained, achieving the effect of accurately identifying changes in land surface features.

[0008] In some embodiments, before acquiring remote sensing images of the target area and multi-round question-and-answer data on surface changes within the target area, the method further includes: forming multiple data pairs from each remote sensing image, the label corresponding to each remote sensing image, and the multi-round question-and-answer data corresponding to each remote sensing image in the time-series remote sensing dataset; identifying the multiple data pairs using a pre-trained language model to obtain multiple segmentation maps and multiple text descriptions; comparing the multiple segmentation maps and multiple text descriptions respectively to obtain multiple similar features and multiple dissimilar features; fusing the multiple segmentation maps, multiple text descriptions, multiple similar features, and multiple dissimilar features to obtain multiple change detection maps and multiple descriptive features; comparing the multiple change detection maps and multiple descriptive features with the label corresponding to each remote sensing image, and adjusting the parameters of the pre-trained language model based on the comparison results to obtain a spatially aware change detection and text description model.

[0009] In the above embodiments, this application obtains a spatial perception change detection and text description model by feature recognition and fusion of remote sensing images and multi-turn question-and-answer data, and adjusting the parameters of the pre-trained language model based on the final result.

[0010] In some embodiments, multiple change detection maps and multiple descriptive features are compared with the labels corresponding to each remote sensing image. Based on the comparison results, the parameters of the pre-trained language model are adjusted to obtain a spatially aware change detection and text description model. This includes: combining the contextual content of multiple change detection maps, multiple descriptive features, and multi-turn question-answering data corresponding to each remote sensing image to predict the probabilities of multiple change detection maps and multiple descriptive features; comparing the probabilities of multiple change detection maps and multiple descriptive features with the labels corresponding to each remote sensing image to obtain comparison results; and adjusting the parameters of the pre-trained language model based on the comparison results to obtain a spatially aware change detection and text description model.

[0011] In the above embodiments, this application compares the probabilities of multiple change detection maps and multiple descriptive features with the labels corresponding to each remote sensing image. This allows for accurate adjustment of the parameters of the pre-trained language model based on the comparison results, resulting in a spatially perceptive change detection and text description model.

[0012] In some embodiments, before combining each remote sensing image, the corresponding label for each remote sensing image, and the multi-turn question-and-answer data corresponding to each remote sensing image into multiple data pairs, the method further includes: cropping high-resolution remote sensing images through a sliding window to obtain a time-series remote sensing dataset; labeling the land surface changes in each remote sensing image in the time-series remote sensing dataset to obtain the corresponding label for each remote sensing image; identifying the land surface change data in the time-series remote sensing dataset, and generating multi-turn question-and-answer data corresponding to each remote sensing image in CoT format after answering the user's questions using the land surface change data.

[0013] In the above embodiments, this application generates multi-round question-and-answer data corresponding to each remote sensing image in CoT format after answering the user's questions using surface change data, so as to facilitate the accurate identification of the question-and-answer data by the subsequent model.

[0014] In some embodiments, a preset spatial perception change detection and text description model is used to identify remote sensing images and multi-round question-and-answer data to obtain a change detection map of the target area and a corresponding text description, including: identifying the surface change portion in the remote sensing images and multi-round question-and-answer data through the spatial perception change detection and text description model to obtain a change detection map of the target area and a corresponding text description.

[0015] In the above embodiments, this application identifies surface changes in remote sensing images and multi-round question-and-answer data by using spatial perception change detection and text description models, which can accurately obtain change detection maps and corresponding text descriptions of the target area.

[0016] In some embodiments, acquiring remote sensing images of a target area and multi-round question-and-answer data on surface changes within the target area includes: capturing remote sensing images of the target area through a sliding window to acquire remote sensing images of the target area; and inputting a user's question to acquire multi-round question-and-answer data.

[0017] In the above embodiments, this application uses a sliding window to capture remote sensing images of the target area to identify the answers to user questions and obtain accurate multi-round question-and-answer data.

[0018] In some embodiments, before acquiring remote sensing images of the target area and multi-round question-and-answer data on surface changes within the target area, the method further includes: acquiring dual-temporal images of the target area; identifying change images in the dual-temporal images to obtain remote sensing images of the target area.

[0019] In the above embodiments, this application can accurately acquire remote sensing images from the changing images in dual-temporal images.

[0020] Secondly, embodiments of this application provide an apparatus for identifying changes in land surface features, comprising:

[0021] The acquisition module is used to acquire remote sensing images of the target area and multi-round question-and-answer data of surface changes within the target area. The multi-round question-and-answer data is generated based on historical actual surface change data within the target area.

[0022] The recognition module is used to recognize remote sensing images and multi-turn question-and-answer data using a preset spatially perceived change detection and text description model, to obtain change detection maps and corresponding text descriptions of target areas. The spatially perceived change detection and text description model is obtained by adjusting the parameters of a pre-trained language model by fusing multiple change detection maps and multiple description features obtained from multiple segmentation maps, multiple text descriptions, multiple similar features, and multiple difference features. Multiple segmentation maps and multiple text descriptions are obtained by recognizing multiple data pairs using the pre-trained language model. Multiple similar features and multiple difference features are obtained by comparing multiple segmentation maps and multiple text descriptions respectively. Multiple data pairs are composed of each remote sensing image in the time-series remote sensing dataset, the label corresponding to each remote sensing image, and the multi-turn question-and-answer data corresponding to each remote sensing image.

[0023] Optionally, the device further includes:

[0024] The training module is used to form multiple data pairs from each remote sensing image, the label corresponding to each remote sensing image, and the multi-round question and answer data corresponding to each remote sensing image in the time-series remote sensing dataset before the acquisition module acquires remote sensing images of the target area and multi-round question and answer data of surface changes in the target area.

[0025] Multiple data pairs are identified using a pre-trained language model to obtain multiple segmentation maps and multiple text descriptions;

[0026] By comparing multiple segmentation images and multiple text descriptions, multiple similar features and multiple dissimilar features are obtained;

[0027] By fusing multiple segmentation maps, multiple text descriptions, multiple similar features, and multiple difference features, multiple change detection maps and multiple descriptive features are obtained;

[0028] Multiple change detection maps and multiple descriptive features are compared with the labels corresponding to each remote sensing image. The parameters of the pre-trained language model are adjusted based on the comparison results to obtain a spatially perceptive change detection and text description model.

[0029] Optionally, the acquisition module is specifically used for:

[0030] By combining multiple change detection maps, multiple descriptive features, and the contextual content of multi-round question-and-answer data corresponding to each remote sensing image, the probability of outputting multiple change detection maps and multiple descriptive features is predicted.

[0031] The probabilities of multiple change detection maps and multiple descriptive features are compared with the labels corresponding to each remote sensing image to obtain the comparison results;

[0032] Based on the comparison results, the parameters of the pre-trained language model were adjusted to obtain a spatially perceptual change detection and text description model.

[0033] Optionally, the device further includes:

[0034] The generation module is used by the training module to extract high-resolution remote sensing images through a sliding window before combining each remote sensing image, the label corresponding to each remote sensing image, and the multi-round question-and-answer data corresponding to each remote sensing image into multiple data pairs, so as to obtain the time-series remote sensing dataset.

[0035] By labeling the surface changes in each remote sensing image in the time-series remote sensing dataset, the label corresponding to each remote sensing image is obtained.

[0036] The system identifies surface change data in time-series remote sensing datasets, and after answering user-provided questions using this data, generates multi-round question-and-answer data for each remote sensing image in CoT format.

[0037] Optionally, the prediction module is specifically used for:

[0038] By using a spatially perceptual change detection and text description model, surface changes in remote sensing images and multi-round question-and-answer data are identified, resulting in a change detection map of the target area and a corresponding text description.

[0039] Optionally, the acquisition module is specifically used for:

[0040] Remote sensing images of the target area are obtained by capturing the target area through a sliding window;

[0041] Input the user's question and obtain multi-round question and answer data.

[0042] Optionally, the device further includes:

[0043] The identification module is used to acquire dual-temporal images of the target area before the acquisition module acquires remote sensing images of the target area and multi-round question-and-answer data on surface changes within the target area;

[0044] By identifying changes in the dual-temporal images, a remote sensing image of the target area can be obtained.

[0045] Thirdly, embodiments of this application provide an electronic device, including a processor and a memory, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the steps of the method provided in the first aspect above are performed.

[0046] Fourthly, embodiments of this application provide a readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the steps of the method provided in the first aspect above.

[0047] Other features and advantages of this application will be set forth in the following description and will be apparent in part from the description or may be learned by practicing embodiments of this application. Attached Figure Description

[0048] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0049] Figure 1 A flowchart illustrating a method for identifying changes in land surface features, provided in an embodiment of this application;

[0050] Figure 2 A flowchart illustrating an implementation method for identifying changes in land surface features, provided as an embodiment of this application;

[0051] Figure 3 A schematic block diagram of a device for identifying changes in land surface features, provided as an embodiment of this application;

[0052] Figure 4This is a schematic block diagram of a device for identifying changes in surface features, provided as an embodiment of this application. Detailed Implementation

[0053] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0054] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0055] First, some of the terms used in the embodiments of this application will be explained to facilitate understanding by those skilled in the art.

[0056] Chain-of-Thought (CoT) format is a technique used to enhance the reasoning capabilities of large language models (LLMs) and explain their thought processes.

[0057] Two-phase images refer to two or more images of the same scene taken at different times. These images are usually taken at different times in order to capture information that changes over time.

[0058] This application is applied to scenarios involving remote sensing data and text recognition. Specifically, it involves generating CoT format question-and-answer data by creating and designing a dual-temporal remote sensing dataset. By combining a pre-trained model and fine-tuning strategies, it achieves accurate identification of spatial changes and context-coherent semantic text descriptions, supporting large-scale surface prediction.

[0059] Change detection (CD) is a core task in remote sensing image analysis, aiming to identify surface changes, such as new buildings and road expansions, by analyzing images from different time periods. Traditional methods mainly rely on pixel-level or feature-level image segmentation techniques combined with optical remote sensing data (such as Landsat and Sentinel-2) for analysis. These methods can effectively capture large-scale surface changes by comparing the radiometric or textural features of the images. High-resolution remote sensing data provides shorter revisit periods and higher spatial resolution, driving research into fine-scale change detection. However, existing remote sensing datasets typically only provide static descriptive statements and lack multi-turn question-and-answer data that supports contextual coherence. Furthermore, traditional methods face multiple challenges in high-resolution imagery scenarios. First, mixed pixels and complex backgrounds increase noise in the segmentation results, making it difficult to accurately distinguish subtle changes.

[0060] To this end, this application acquires remote sensing images of the target area and multi-turn question-and-answer data on surface changes within the target area. The multi-turn question-and-answer data is generated based on historical actual surface change data within the target area. A pre-set spatially-aware change detection and text description model is used to identify the remote sensing images and multi-turn question-and-answer data, resulting in change detection maps and corresponding text descriptions for the target area. The spatially-aware change detection and text description model is obtained by fusing multiple segmentation maps, multiple text descriptions, multiple similar features, and multiple dissimilar features to adjust the parameters of a pre-trained language model. Multiple segmentation maps and multiple text descriptions are obtained by identifying multiple data pairs using the pre-trained language model. Multiple similar features and multiple dissimilar features are obtained by comparing multiple segmentation maps and multiple text descriptions separately. Multiple data pairs are composed of each remote sensing image in the time-series remote sensing dataset, the corresponding label for each remote sensing image, and the multi-turn question-and-answer data corresponding to each remote sensing image. By using the pre-trained spatially-aware change detection and text description model to identify remote sensing images and multi-turn question-and-answer data, and by combining the identification results of different data, more accurate surface change prediction results can be obtained, achieving the effect of accurately identifying changes in surface features.

[0061] In this embodiment of the application, the executing entity can be the device for identifying changes in land features in the system for identifying changes in land features. In practical applications, the device for identifying changes in land features can be electronic devices such as terminal devices and servers, and there are no restrictions on this.

[0062] The following is combined Figure 1 The method for identifying changes in land surface features according to embodiments of this application will be described in detail.

[0063] Please refer to Figure 1 , Figure 1A flowchart illustrating a method for identifying changes in land surface features provided in this application embodiment is shown below. Figure 1 The methods shown for identifying changes in land surface features include:

[0064] Step 110: Acquire remote sensing images of the target area and multi-round question-and-answer data on surface changes within the target area.

[0065] The multi-turn question-and-answer data is generated based on historical actual surface change data within the target area, or it can be obtained through multi-turn dialogues between the question-and-answer model in the robot and the user. The target area can be any surface region, including mountains, farms, urban buildings, and rural roads.

[0066] In some embodiments of this application, before acquiring remote sensing images of the target area and multi-round question-and-answer data on surface changes within the target area, Figure 1 The method further includes: forming multiple data pairs from each remote sensing image in the time-series remote sensing dataset, the corresponding label for each remote sensing image, and the multi-turn question-and-answer data for each remote sensing image; identifying the multiple data pairs using a pre-trained language model to obtain multiple segmentation maps and multiple text descriptions; comparing the multiple segmentation maps and multiple text descriptions to obtain multiple similar features and multiple dissimilar features; fusing the multiple segmentation maps, multiple text descriptions, multiple similar features, and multiple dissimilar features to obtain multiple change detection maps and multiple descriptive features; comparing the multiple change detection maps and multiple descriptive features with the label corresponding to each remote sensing image, and adjusting the parameters of the pre-trained language model based on the comparison results to obtain a spatially aware change detection and text description model.

[0067] In the aforementioned process, this application utilizes feature recognition and fusion of remote sensing images and multi-round question-and-answer data. Based on the final results, the parameters of the pre-trained language model can be adjusted to obtain a spatially perceptual change detection and text description model. The comparison results include differences and loss calculations.

[0068] The time-series remote sensing dataset includes multiple remote sensing images of the target area taken by different satellites. The label for each remote sensing image was obtained based on annotations by relevant personnel. Each dataset includes the remote sensing image, its corresponding label, and multi-turn question-and-answer data. Similarity features are obtained through similarity calculations, while the remaining features are difference features.

[0069] Optionally, before combining each remote sensing image, the corresponding label, and the multi-turn question-and-answer data in the time-series remote sensing dataset into multiple data pairs, preprocessing of the high-resolution time-series remote sensing dataset is also included to obtain each remote sensing image, the corresponding label, and the multi-turn question-and-answer data.

[0070] Specifically: First, this method is based on an original, constructed high-resolution temporal remote sensing dataset, containing bi-temporal images (A / B phases), image labels (annotating changed regions and categories), and question-and-answer data in CoT format. The bi-temporal images are then normalized using the following formula:

[0071] ;

[0072] in, Let I' be the original image value of channel c at position (h,w), where μc and σc are the mean and standard deviation of channel c, respectively, and I′(c,h,w) is the normalized image value. Then, the segmentation label pixel values ​​are mapped to categories, which are divided into buildings, roads, and unchanged parts. The calculation formula is as follows:

[0073] ;

[0074] in, The values ​​are the original pixel values, and L(h,w) is the category label. Category label values ​​of 2, 1, and 0 represent buildings, roads, and unchanged parts, respectively. Equation Chapter (Next) Section 1

[0075] Optionally, the labels for each remote sensing image are obtained as follows: The sample annotations consist of two main parts. The first part is a mask image of the changed area location, which is manually and visually meticulously annotated. The second part is a textual CoT descriptive annotation. First, CoT format question-and-answer data is created for the dual-temporal images, containing multi-turn question-and-answer pairs (e.g., "Are there any changes between these two images?" "Yes." "What type of new changed area is it?" "A new building has been added." "Where is the building?" "In the middle of the image."). This ensures the question-and-answer format is coherent. Then, the CoT question-and-answer data is encoded using the pre-trained language model BERT to generate a token word sequence. The calculation formula is as follows:

[0076] ;

[0077] Among them, c j For the j-th answer, It is a sequence of lexical elements and , This represents the maximum sequence length. The effective length helps the model focus only on meaningful words when processing sequence data, ignoring padding, thereby improving the efficiency and accuracy of model training and prediction. The formula for calculating the effective length is as follows:

[0078] ;

[0079] in, Let be the effective length of the j-th answer, and pad be the padding marker. Then, CoT question-answering is paired with bi-temporal images to ensure semantic consistency between the question and the changing region. Finally, the dataset is divided into training, validation and test sets for subsequent model training, optimization and testing.

[0080] Optionally, multiple data pairs are identified using a pre-trained language model to obtain multiple segmentation images and multiple text descriptions; by comparing the multiple segmentation images and multiple text descriptions, multiple similar features and multiple dissimilar features are obtained, including:

[0081] The system identifies changed regions in dual-temporal images and generates segmentation maps. The Change Caption (CC) branch focuses on generating semantic text descriptions related to the changed parts, supporting a CoT question-and-answer format. First, it uses dual-temporal images... As input, a pre-trained visual model is used to extract multi-scale features, calculated using the following formula:

[0082] ;

[0083] in, Temporal images Extracted multi-scale features and N is the number of feature layers.

[0084] Optionally, multiple segmentation maps, multiple text descriptions, multiple similarity features, and multiple difference features are fused to obtain multiple change detection maps and multiple descriptive features, including: features The CoT problem P={p1,p2,…} and the answer C={c1,c2,…} are used as inputs. In the change detection branch, feature differences and similarities are calculated. Feature differences highlight pixel-level changes between two temporal images, and cosine similarity measures the degree of similarity between feature vectors. The combination of these two factors provides comprehensive information for subsequent Transformer fusion, helping the model to accurately locate change regions. The calculation formula is as follows:

[0085] ;

[0086] Where Di represents the difference features and Si represents the cosine similarity, the Transformer is used to fuse the features and project them onto the segmentation space, and the calculation formula is as follows:

[0087] :

[0088] in, These are the fused features of the i-th layer. The Transformer models Di and Si through a multi-head self-attention mechanism to capture the changing spatiotemporal dependencies, and then combines the multi-scale features. splicing and merging through convolution operations into This generates a predicted segmentation map, calculated using the following formula:

[0089] :

[0090] in, To predict the segmentation map, where K is the number of classes, Conv uses a 1×1 convolution kernel to... Mapping to the class space, Upsample restores the feature map to its original resolution using bilinear interpolation.

[0091] Encode the CoT question and answer in the change description branch, and calculate the formula as follows:

[0092] :

[0093] The encoded result contains semantic information of the text, which is used for cross-modal feature projection:

[0094] (10)Equation Chapter (Next) Section 1

[0095] (11)Equation Chapter (Next) Section 1

[0096] :

[0097] in, , , As a projection layer, it outputs descriptive features. It preserves contextual information and outputs a segmentation map. and descriptive features .

[0098] Optionally, multiple change detection maps and multiple descriptive features are compared with the labels corresponding to each remote sensing image. Based on the comparison results, the parameters of the pre-trained language model are adjusted to obtain a spatially aware change detection and text description model, including:

[0099] In the sequence generation stage, to describe features and historical issues As input, a sequence is generated using the Transformer decoder, and combined with historical context, the calculation formula is as follows:

[0100] ;

[0101] in, V is the vocabulary size, and the text is generated through beam search. The calculation formula is as follows:

[0102] ;

[0103] Output prediction probability (Training) or text sequence (Inference) During the training phase, the loss function is defined as follows: The change detection loss is as follows:

[0104] ;

[0105] Where C is the number of categories, and H and W are the segmentation map dimensions. To predict probabilities, This is a real label.

[0106] The change describes the loss as follows:

[0107] ;

[0108] Where L is the sequence length and V is the vocabulary size. For the real token at position l, To predict probabilities.

[0109] The joint loss formula is shown in Equation Chapter (Next) Section 1

[0110] ;

[0111] Using the Adam optimizer, the learning rate The gradient accumulation steps are as follows:

[0112] ;

[0113] Where B is the batch size, used for initial training and joint training. Future plans include expanding the dataset to hundreds of images, adding different forms of CoT question-and-answer data, and optimizing model performance.

[0114] In some embodiments of this application, multiple change detection maps and multiple descriptive features are compared with the labels corresponding to each remote sensing image. Based on the comparison results, the parameters of the pre-trained language model are adjusted to obtain a spatially aware change detection and text description model. This includes: combining the contextual content of multiple change detection maps, multiple descriptive features, and multi-turn question-answering data corresponding to each remote sensing image to predict the probabilities of multiple change detection maps and multiple descriptive features; comparing the probabilities of multiple change detection maps and multiple descriptive features with the labels corresponding to each remote sensing image to obtain comparison results; and adjusting the parameters of the pre-trained language model based on the comparison results to obtain a spatially aware change detection and text description model.

[0115] In the above process, this application compares the probabilities of multiple change detection maps and multiple descriptive features with the labels corresponding to each remote sensing image. This allows for accurate adjustment of the parameters of the pre-trained language model based on the comparison results, resulting in a spatially perceptive change detection and text description model.

[0116] In some embodiments of this application, before combining each remote sensing image, the corresponding label for each remote sensing image, and the multi-turn question-and-answer data corresponding to each remote sensing image into multiple data pairs in the time-series remote sensing dataset, Figure 1 The method also includes: cropping high-resolution remote sensing images through a sliding window to obtain a time-series remote sensing dataset; labeling the land surface changes in each remote sensing image in the time-series remote sensing dataset to obtain a label corresponding to each remote sensing image; identifying land surface change data in the time-series remote sensing dataset; and generating multi-round question-and-answer data corresponding to each remote sensing image in CoT format after answering user-provided questions using the land surface change data.

[0117] In the above process, this application answers the questions provided by users using land surface change data and generates multi-round question-and-answer data for each remote sensing image in CoT format to facilitate accurate identification of the question-and-answer data by the subsequent model.

[0118] The sliding window can be customized as needed. Change data can be obtained by comparing the current surface image with historical surface images.

[0119] In some embodiments of this application, acquiring remote sensing images of a target area and multi-round question-and-answer data on surface changes within the target area includes: capturing remote sensing images of the target area through a sliding window to acquire remote sensing images of the target area; and inputting a user's question to acquire multi-round question-and-answer data.

[0120] In the above process, this application can identify the answers to user questions and obtain accurate multi-round question-and-answer data by using a sliding window to capture remote sensing images of the target area.

[0121] In some embodiments of this application, before acquiring remote sensing images of the target area and multi-round question-and-answer data on surface changes within the target area, Figure 1 The method also includes: acquiring dual-temporal images of the target area; identifying the changing images in the dual-temporal images to obtain a remote sensing image of the target area.

[0122] In the above process, this application can accurately acquire remote sensing images from the changing images in dual-temporal images.

[0123] In the above-mentioned identification of images, videos, and changed images, the image recognition model can be trained based on existing models or historical image data to obtain the data information in the image.

[0124] Step 120: Recognize remote sensing images and multi-round question-and-answer data using a preset spatial perception change detection and text description model to obtain a change detection map of the target area and the corresponding text description.

[0125] Among them, the spatial perception change detection and text description model is obtained by adjusting the parameters of the pre-trained language model by fusing multiple change detection maps and multiple description features obtained by fusing multiple segmentation maps, multiple text descriptions, multiple similar features and multiple difference features. Multiple segmentation maps and multiple text descriptions are obtained by the pre-trained language model to identify multiple data pairs. Multiple similar features and multiple difference features are obtained by comparing multiple segmentation maps and multiple text descriptions respectively. Multiple data pairs are composed of each remote sensing image in the time-series remote sensing dataset, the label corresponding to each remote sensing image, and the multi-turn question-and-answer data corresponding to each remote sensing image.

[0126] In some embodiments of this application, a preset spatial perception change detection and text description model is used to identify remote sensing images and multi-round question-and-answer data to obtain a change detection map of the target area and a corresponding text description. This includes: identifying the surface change portion in the remote sensing images and multi-round question-and-answer data using the spatial perception change detection and text description model to obtain a change detection map of the target area and a corresponding text description.

[0127] In the above process, this application identifies the surface changes in remote sensing images and multi-round question-and-answer data through spatial perception change detection and text description models, which can accurately obtain the change detection map and corresponding text description of the target area.

[0128] In the above Figure 1In the process shown, this application acquires remote sensing images of the target area and multi-round question-and-answer data on surface changes within the target area. The multi-round question-and-answer data is generated based on historical actual surface change data within the target area. A preset spatially perceived change detection and text description model is used to identify the remote sensing images and multi-round question-and-answer data to obtain change detection maps and corresponding text descriptions of the target area. The spatially perceived change detection and text description model is obtained by adjusting the parameters of a pre-trained language model by fusing multiple change detection maps and multiple description features obtained from multiple segmentation maps, multiple text descriptions, multiple similar features, and multiple difference features. Multiple segmentation maps and multiple text descriptions are obtained by identifying multiple data pairs through the pre-trained language model. Multiple similar features and multiple difference features are obtained by comparing multiple segmentation maps and multiple text descriptions respectively. Multiple data pairs are composed of each remote sensing image in the time-series remote sensing dataset, the label corresponding to each remote sensing image, and the multi-round question-and-answer data corresponding to each remote sensing image. By using a pre-trained spatial perception change detection and text description model to identify remote sensing images and multi-turn question-and-answer sessions, more accurate land surface change prediction results can be obtained by combining the identification results of different data, thus achieving the effect of accurately identifying changes in land surface features.

[0129] The following is combined Figure 2 The implementation method for identifying changes in land surface features according to embodiments of this application will be described in detail.

[0130] Please refer to Figure 2 , Figure 2 A flowchart illustrating an implementation method for identifying changes in land surface features, as provided in this application embodiment, is shown below. Figure 2 The implementation methods for identifying changes in land surface features shown include:

[0131] Step 210: Image recognition.

[0132] Specifically: Identify remote sensing images of the target area and obtain image features.

[0133] Step 220: Question and Answer Recognition.

[0134] Specifically: Identify the multi-round question-and-answer data corresponding to each remote sensing image to obtain text features.

[0135] Step 230: Feature fusion.

[0136] Specifically, by fusing image features and text features, as well as similarity and difference features corresponding to segmentation maps and text descriptions, multiple change detection maps and multiple descriptive features are obtained.

[0137] Step 240: Surface identification.

[0138] Specifically, multiple change detection maps and multiple descriptive features are transformed into prediction maps and text descriptions to obtain recognition results.

[0139] also, Figure 2 The specific methods and steps shown can be found in [reference]. Figure 1 The method shown will not be elaborated further here.

[0140] The previous text passed Figures 1-2 The method for identifying changes in land surface features is described below, in conjunction with... Figures 3-4 Describes a device for identifying changes in surface features.

[0141] Please refer to Figure 3 This is a schematic block diagram of a device 300 for identifying changes in land surface features provided in an embodiment of this application. The device 300 can be a module, program segment, or code on an electronic device. This device 300 is similar to the one described above. Figure 1 The method implementation corresponds to this and can be executed. Figure 1 The various steps involved in the method embodiments and the specific functions of the device 300 can be found in the following description. To avoid repetition, detailed descriptions are appropriately omitted here.

[0142] Optionally, the device 300 includes:

[0143] The acquisition module 310 is used to acquire remote sensing images of the target area and multi-round question-and-answer data of surface changes in the target area. The multi-round question-and-answer data is generated based on historical actual surface change data in the target area.

[0144] The recognition module 320 is used to recognize remote sensing images and multi-turn question-and-answer data through a preset spatially perceived change detection and text description model to obtain a change detection map and corresponding text description of the target area. The spatially perceived change detection and text description model is obtained by adjusting the parameters of a pre-trained language model by fusing multiple change detection maps and multiple description features obtained by fusing multiple segmentation maps, multiple text descriptions, multiple similar features and multiple difference features. Multiple segmentation maps and multiple text descriptions are obtained by recognizing multiple data pairs through a pre-trained language model. Multiple similar features and multiple difference features are obtained by comparing multiple segmentation maps and multiple text descriptions respectively. Multiple data pairs are composed of each remote sensing image in the time-series remote sensing dataset, the label corresponding to each remote sensing image, and the multi-turn question-and-answer data corresponding to each remote sensing image.

[0145] Optionally, the device further includes:

[0146] The training module, before the acquisition module acquires remote sensing images of the target area and multi-round question-and-answer data on surface changes within the target area, assembles each remote sensing image, its corresponding label, and its corresponding multi-round question-and-answer data from the time-series remote sensing dataset into multiple data pairs. A pre-trained language model is used to identify these data pairs, obtaining multiple segmentation maps and multiple text descriptions. The multiple segmentation maps and multiple text descriptions are compared to obtain multiple similar features and multiple dissimilar features. These are then fused to obtain multiple change detection maps and multiple descriptive features. Finally, the multiple change detection maps and multiple descriptive features are compared with the labels corresponding to each remote sensing image, and the parameters of the pre-trained language model are adjusted based on the comparison results to obtain a spatially aware change detection and text description model.

[0147] Optionally, the acquisition module is specifically used for:

[0148] By combining multiple change detection maps, multiple descriptive features, and the contextual content of multi-round question-and-answer data corresponding to each remote sensing image, the probabilities of multiple change detection maps and multiple descriptive features are predicted and output. The probabilities of multiple change detection maps and multiple descriptive features are compared with the labels corresponding to each remote sensing image to obtain the comparison results. Based on the comparison results, the parameters of the pre-trained language model are adjusted to obtain a spatially aware change detection and text description model.

[0149] Optionally, the device further includes:

[0150] The generation module is used by the training module to extract high-resolution remote sensing images through a sliding window to obtain a time-series remote sensing dataset before combining each remote sensing image, its corresponding label, and its corresponding multi-turn question-and-answer data into multiple data pairs. It then labels each remote sensing image by annotating the land surface changes in each image; identifies the land surface change data in the time-series remote sensing dataset; and generates multi-turn question-and-answer data for each remote sensing image in CoT format after answering user-provided questions using the land surface change data.

[0151] Optionally, the prediction module is specifically used for:

[0152] By using a spatially perceptual change detection and text description model, surface changes in remote sensing images and multi-round question-and-answer data are identified, resulting in a change detection map of the target area and a corresponding text description.

[0153] Optionally, the acquisition module is specifically used for:

[0154] Remote sensing images of the target area are captured by sliding a window to obtain the remote sensing image of the target area; multiple rounds of question and answer data are obtained by inputting the user's question.

[0155] Optionally, the device further includes:

[0156] The identification module is used to acquire dual-temporal images of the target area before the acquisition module acquires remote sensing images of the target area and multi-round question-and-answer data on surface changes within the target area; and to identify the changing images in the dual-temporal images to obtain remote sensing images of the target area.

[0157] Please refer to Figure 4 This is a schematic block diagram of a device for identifying changes in land surface features according to an embodiment of this application. The device may include a memory 410 and a processor 420. Optionally, the device may further include a communication interface 430 and a communication bus 440. This device is similar to the one described above. Figure 1 The method implementation corresponds to this and can be executed. Figure 1 The specific functions of the device involved in the method embodiments can be found in the following description.

[0158] Specifically, memory 410 is used to store computer-readable instructions.

[0159] Processor 420 is used to process readable instructions stored in memory and is capable of executing... Figure 1 Each step in the method.

[0160] The communication interface 430 is used for signaling or data communication with other node devices. For example, it is used for communication with a server or terminal, or for communication with other device nodes, but the embodiments of this application are not limited thereto.

[0161] Communication bus 440 is used to enable direct communication between the above components.

[0162] In this embodiment, the communication interface 430 of the device is used for signaling or data communication with other node devices. The memory 410 can be high-speed RAM or non-volatile memory, such as at least one disk storage device. Optionally, the memory 410 can also be at least one storage device located remotely from the aforementioned processor. The memory 410 stores computer-readable instructions, which, when executed by the processor 420, enable the electronic device to perform the aforementioned... Figure 1The method process is shown. Processor 420 can be used on device 300 and is used to perform the functions in this application. Exemplarily, the processor 420 described above can be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, and the embodiments of this application are not limited thereto.

[0163] This application embodiment also provides a readable storage medium, wherein when the computer program is executed by a processor, it performs the following... Figure 1 The method process executed by the electronic device in the illustrated method embodiment.

[0164] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the device described above can be referred to the corresponding process in the aforementioned method, and will not be elaborated further here.

[0165] In summary, this application provides a method, apparatus, device, and storage medium for identifying changes in land surface features. The method includes acquiring remote sensing images of a target area and multi-turn question-and-answer data on land surface changes within the target area. The multi-turn question-and-answer data is generated based on historical actual land surface change data within the target area. A preset spatially-aware change detection and text description model is used to identify the remote sensing images and multi-turn question-and-answer data, resulting in a change detection map and corresponding text description for the target area. The spatially-aware change detection and text description model is obtained by adjusting the parameters of a pre-trained language model using multiple change detection maps and multiple descriptive features obtained by fusing multiple segmentation maps, multiple text descriptions, multiple similar features, and multiple difference features. The multiple segmentation maps and multiple text descriptions are obtained by identifying multiple data pairs using the pre-trained language model. The multiple similar features and multiple difference features are obtained by comparing multiple segmentation maps and multiple text descriptions respectively. The multiple data pairs are composed of each remote sensing image in a time-series remote sensing dataset, the label corresponding to each remote sensing image, and the multi-turn question-and-answer data corresponding to each remote sensing image. This method can achieve accurate identification of changes in land surface features.

[0166] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0167] In addition, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0168] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0169] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application. It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0170] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

[0171] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

Claims

1. A method of identifying a change in a surface feature, the method comprising: The method comprises: acquiring remote sensing images of a target area and multi-turn question and answer data of ground surface changes in the target area, wherein the multi-turn question and answer data is generated according to historical actual ground surface change data in the target area; identifying the remote sensing images and the multi-turn question and answer data through a preset spatially-aware change detection and text description model to obtain a change detection map of the target area and corresponding text description, wherein the spatially-aware change detection and text description model is obtained by adjusting parameters of a pre-trained language model through a plurality of change detection maps and a plurality of description features obtained by fusing a plurality of segmentation maps, a plurality of text descriptions, a plurality of similar features and a plurality of difference features, the plurality of segmentation maps and the plurality of text descriptions are obtained by identifying a plurality of data pairs through the pre-trained language model, the plurality of similar features and the plurality of difference features are obtained by comparing the plurality of segmentation maps and the plurality of text descriptions respectively, and the plurality of data pairs are composed of each remote sensing image in a time-series remote sensing dataset, a label corresponding to each remote sensing image and multi-turn question and answer data corresponding to each remote sensing image.

2. The method of claim 1, wherein, Before the acquiring remote sensing images of a target area and multi-turn question and answer data of ground surface changes in the target area, the method further comprises: composing each remote sensing image in the time-series remote sensing dataset, a label corresponding to each remote sensing image and multi-turn question and answer data corresponding to each remote sensing image into the plurality of data pairs; identifying the plurality of data pairs through the pre-trained language model to obtain the plurality of segmentation maps and the plurality of text descriptions; comparing the plurality of segmentation maps and the plurality of text descriptions respectively to obtain the plurality of similar features and the plurality of difference features; fusing the plurality of segmentation maps, the plurality of text descriptions, the plurality of similar features and the plurality of difference features to obtain the plurality of change detection maps and the plurality of description features; comparing the plurality of change detection maps and the plurality of description features with the label corresponding to each remote sensing image, and adjusting parameters of the pre-trained language model according to the obtained comparison result to obtain the spatially-aware change detection and text description model.

3. The method of claim 2, wherein, The comparing the plurality of change detection maps and the plurality of description features with the label corresponding to each remote sensing image, and adjusting parameters of the pre-trained language model according to the obtained comparison result to obtain the spatially-aware change detection and text description model comprises: predicting probabilities of the plurality of change detection maps and the plurality of description features in combination with the plurality of change detection maps, the plurality of description features and context content of the multi-turn question and answer data corresponding to each remote sensing image; comparing the probabilities of the plurality of change detection maps and the plurality of description features with the label corresponding to each remote sensing image respectively to obtain a comparison result; adjusting parameters of the pre-trained language model according to the comparison result to obtain the spatially-aware change detection and text description model.

4. The method of claim 2, wherein, Before the forming of the plurality of data pairs by each remote sensing image in the time-series remote sensing dataset, a corresponding label of each remote sensing image, and a corresponding multi-turn question-answering data of each remote sensing image, the method further comprises: The time-series remote sensing dataset is obtained by intercepting the high-resolution remote sensing image through a sliding window. The corresponding label of each remote sensing image is obtained by labeling the ground surface change in each remote sensing image in the time-series remote sensing dataset. The corresponding multi-turn question-answering data of each remote sensing image is obtained by identifying the ground surface change data in the time-series remote sensing dataset and answering the question provided by the user in the CoT format.

5. The method according to any one of claims 1 to 3, characterized in that, The change detection graph and the corresponding text description of the target area are obtained by identifying the remote sensing image and the multi-turn question-answering data through the preset spatially-aware change detection and text description model. The change detection graph and the corresponding text description of the target area are obtained by identifying the ground surface change part in the remote sensing image and the multi-turn question-answering data through the spatially-aware change detection and text description model.

6. The method according to any one of claims 1 to 3, characterized in that, The remote sensing image of the target area and the multi-turn question-answering data of the ground surface change in the target area are obtained by: The remote sensing image of the target area is obtained by intercepting the remote sensing image of the target area through a sliding window. The multi-turn question-answering data is obtained by inputting the user's question.

7. The method according to any one of claims 1 to 3, characterized in that, Before the remote sensing image of the target area and the multi-turn question-answering data of the ground surface change in the target area are obtained, the method further comprises: The double-time-phase image of the target area is obtained. The remote sensing image of the target area is obtained by identifying the change image in the double-time-phase image.

8. An apparatus for identifying changes in surface features, the apparatus comprising: Comprise: The acquisition module is configured to obtain the remote sensing image of the target area and the multi-turn question-answering data of the ground surface change in the target area, wherein the multi-turn question-answering data is generated according to the historical actual ground surface change data in the target area. The identification module is configured to identify the remote sensing image and the multi-turn question-answering data through a preset spatially-aware change detection and text description model to obtain a change detection graph and a corresponding text description of the target area, wherein the spatially-aware change detection and text description model is obtained by adjusting the parameters of a pre-trained language model through a plurality of segmentation graphs, a plurality of text descriptions, a plurality of similar features, and a plurality of difference features, the plurality of segmentation graphs and the plurality of text descriptions are obtained by identifying a plurality of data pairs through the pre-trained language model, the plurality of similar features and the plurality of difference features are obtained by comparing the plurality of segmentation graphs and the plurality of text descriptions, respectively, and the plurality of data pairs are formed by each remote sensing image in a time-series remote sensing dataset, a corresponding label of each remote sensing image, and a corresponding multi-turn question-answering data of each remote sensing image.

9. An electronic device, comprising: Comprise: The memory stores computer readable instructions, and when the computer readable instructions are executed by the processor, the steps in the method of any one of claims 1-7 are run.

10. A computer-readable storage medium, characterized in that, Comprise: Computer program, which when run on a computer, causes the computer to perform the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Remote sensing cultivated land change detection method and system considering phenological characteristics

    CN111768101A

  • Building target damage assessment method based on multi-modal remote sensing data

    CN119810523A