Large model remote sensing small sample change detection method and system combined with graph attention

By performing radiation correction, atmospheric correction and geographic registration on remote sensing images and combining them with the graph attention feature encoder method, the problems of low accuracy and poor generalization ability of traditional remote sensing change detection under small sample data are solved, and efficient and intelligent remote sensing change detection is achieved.

CN120707799AActive Publication Date: 2025-09-26ZHONGKAN MAIPU (JIANGSU) TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511098247.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-09-26
Estimated Expiration
2045-08-06

AI Technical Summary

Technical Problem

Traditional remote sensing change detection methods suffer from low accuracy and poor generalization ability when processing small sample data.

Method used

A large-scale remote sensing small-sample change detection method combined with graph attention is adopted. By performing radiation correction and atmospheric correction on the original remote sensing images, standardized remote sensing images are generated. The images of different phases are aligned to a unified geographic coordinate system through geo-registration. The graph attention feature encoder is used to capture the spatiotemporal topological relationship of the ground objects, and a target remote sensing small-sample change detection model is constructed to output a change type map.

Benefits of technology

It improves the accuracy of change detection in small sample scenarios, solves the problem of traditional methods relying on large amounts of labeled data, and realizes efficient and intelligent remote sensing image analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707799A_ABST
    Figure CN120707799A_ABST
Patent Text Reader

Abstract

The invention discloses a large model remote sensing small sample change detection method and system combined with graph attention, and relates to the technical field of artificial intelligence, and the method comprises the steps: firstly carrying out the radiation correction and atmospheric correction of an original remote sensing image, and generating a standardized remote sensing image; the standardized images of different time phases are aligned to a unified geographic coordinate system through geographic registration; obtaining the aligned reference time phase and change time phase images; and inputting the two time phase images into a pre-trained target remote sensing small sample change detection model containing a graph attention feature encoder, and outputting a change type graph of the target area. According to the method, the ground feature space-time topological relation is captured through a graph attention mechanism, the problem that a traditional method depends on a large amount of labeled data is solved, and the change detection accuracy in a small sample scene is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a large-model remote sensing small sample change detection method and system combined with graph attention. Background Art

[0002] With the development of remote sensing technology, change detection has been widely used in fields such as resource monitoring and urban planning. Traditional remote sensing change detection methods suffer from low accuracy and poor generalization when processing small sample data. Summary of the Invention

[0003] The purpose of the present invention is to provide a large-model remote sensing small sample change detection method and system combined with graph attention.

[0004] In a first aspect, an embodiment of the present invention provides a large-model remote sensing small-sample change detection method combined with graph attention, comprising:

[0005] Perform radiation correction and atmospheric correction on the original remote sensing image data of the target area to generate standardized remote sensing images;

[0006] Through georeferencing, standardized remote sensing images of different phases are aligned to a unified geographic coordinate system;

[0007] Acquire aligned reference time phase remote sensing images and change time phase remote sensing images; the change time phase remote sensing images are remote sensing images acquired at a preset acquisition time interval of the reference time phase remote sensing images;

[0008] The reference time phase remote sensing image and the change time phase remote sensing image are input into a pre-trained target remote sensing small sample change detection model to obtain a change type map of the target area.

[0009] In one possible implementation, the target remote sensing small sample change detection model is trained by the following method, including:

[0010] Obtaining an initial remote sensing small sample change detection model, the initial remote sensing small sample change detection model comprising: a baseline temporal phase feature extraction model and a change discrimination model, the baseline temporal phase feature extraction model and the change discrimination model sharing a same graph attention feature encoder, the baseline temporal phase feature extraction model comprising the graph attention feature encoder and a baseline temporal phase ground feature discriminator, and the change discrimination model comprising the graph attention feature encoder and a change type discriminator;

[0011] The benchmark temporal phase feature extraction model is trained using a benchmark temporal phase remote sensing data instance, and model parameters of the image attention feature encoder and the benchmark temporal phase ground object discriminator included in the benchmark temporal phase feature extraction model are updated to obtain an image attention feature encoder and a benchmark temporal phase ground object discriminator that have completed initial training; wherein the benchmark temporal phase remote sensing data instance includes benchmark temporal phase remote sensing image data;

[0012] Based on the change type discriminator and the graph attention feature encoder that has completed initial training, constructing an initialized change discrimination model;

[0013] The change type discriminator in the initialized change discrimination model is trained using a change phase remote sensing data instance, freezing the model parameters of the graph attention feature encoder that has completed the initial training, and updating the model parameters of the change type discriminator to obtain a change type discriminator that has completed the initial training; wherein the change phase remote sensing data instance includes change phase remote sensing image data;

[0014] Constructing a change discrimination model that has completed initial training based on the graph attention feature encoder that has completed initial training and the change type discriminator that has completed initial training;

[0015] The change discrimination model that has completed the initial training is trained using the change phase remote sensing data instance, and the model parameters of the graph attention feature encoder that has completed the initial training and the change type discriminator that has completed the initial training are updated to obtain the graph attention feature encoder that has completed the advanced training and the change type discriminator that has completed the advanced training;

[0016] Based on the graph attention feature encoder that has completed advanced training, the baseline temporal feature discriminator that has completed initial training, and the change type discriminator that has completed advanced training, an intermediate remote sensing small sample change detection model is constructed;

[0017] The intermediate remote sensing small sample change detection model is trained using a small sample change detection data instance to obtain a target remote sensing small sample change detection model; wherein the output of the target remote sensing small sample change detection model is a change type map, and the small sample change detection data instance includes at least one of the following: the baseline phase remote sensing image data, the change phase remote sensing image data.

[0018] In a possible implementation, the reference time phase remote sensing data instance is multispectral sequence data, and the reference time phase remote sensing data instance includes at least one reference time phase multispectral sequence sample data; the change time phase remote sensing data instance is change single frame data, and the change time phase remote sensing data instance includes at least one change time phase hyperspectral single frame sample data;

[0019] The method of training the benchmark temporal phase feature extraction model using a benchmark temporal phase remote sensing data example, updating the model parameters of the graph attention feature encoder and the benchmark temporal phase ground object discriminator contained in the benchmark temporal phase feature extraction model, and obtaining the graph attention feature encoder and the benchmark temporal phase ground object discriminator that have completed initial training, includes:

[0020] For each reference phase multispectral sequence sample data, a plurality of sequence time node images corresponding to the reference phase multispectral sequence sample data are obtained according to a plurality of extracted time sequence nodes;

[0021] Using a plurality of sequence time node images corresponding to the reference time phase multispectral sequence sample data, the reference time phase feature extraction model is trained, and the model parameters of the image attention feature encoder and the reference time phase ground object discriminator included in the reference time phase feature extraction model are updated to obtain the image attention feature encoder and the reference time phase ground object discriminator that have completed initial training;

[0022] The method of using the change phase remote sensing data instance to train the change type discriminator in the initialized change discrimination model, freezing the model parameters of the graph attention feature encoder that has completed the initial training, and updating the model parameters of the change type discriminator to obtain the change type discriminator that has completed the initial training, includes:

[0023] For each hyperspectral single-frame sample data of a changing time phase, data enhancement is performed on the hyperspectral single-frame sample data of the changing time phase to obtain a plurality of enhanced sample data corresponding to the hyperspectral single-frame sample data of the changing time phase;

[0024] The change type discriminator is trained using multiple enhanced sample data corresponding to the single-frame hyperspectral sample data of the change phase, the model parameters of the graph attention feature encoder that has completed the initial training are frozen, and the model parameters of the change type discriminator are updated to obtain the change type discriminator that has completed the initial training.

[0025] In a possible implementation manner, the instance of the time-varying remote sensing data is a time-varying single-frame data, and the time-varying remote sensing data instance includes at least one time-varying hyperspectral single-frame sample data;

[0026] The method of using the change phase remote sensing data instance to train the change type discriminator in the initialized change discrimination model, freezing the model parameters of the graph attention feature encoder that has completed the initial training, and updating the model parameters of the change type discriminator to obtain the change type discriminator that has completed the initial training, includes:

[0027] Cutting a plurality of different image regions from the same single-frame hyperspectral sample data of the changing time phase to obtain a plurality of images to be processed; wherein there are partially overlapping image regions between two of the images to be processed from the same single-frame hyperspectral sample data of the changing time phase;

[0028] determining the plurality of to-be-processed images as sample instances of the initialized change discrimination model;

[0029] Extracting a remote sensing feature representation of the sample instance through the graph attention feature encoder that has completed initial training;

[0030] Determining a change type detection result of the sample instance according to the remote sensing feature representation of the sample instance by the change type discriminator;

[0031] According to the change type detection result and the change type target value of the sample instance, an error parameter is determined, and based on the error parameter, the model parameter of the change type discriminator is updated to obtain the change type discriminator that has completed the initial training.

[0032] In a possible implementation, the step of cropping a plurality of different image regions from the same single-frame hyperspectral sample data of the changing time phase to obtain a plurality of images to be processed includes:

[0033] Determine the size of the image cropping window;

[0034] Sliding the image cropping window to a plurality of different sampling points of the hyperspectral single sample data of the changing time phase, cropping the image areas within the image cropping window respectively, to obtain the plurality of images to be processed;

[0035] The method of cutting out a plurality of different image regions from the same single-frame hyperspectral sample data of the changing time phase to obtain a plurality of images to be processed further includes:

[0036] Determining the size of the image cropping window and its spatial layout position in the hyperspectral single sample data of the changing time phase;

[0037] Determining a plurality of different scale factors corresponding to the single-frame hyperspectral sample data of the changing time phase within a scale transformation range corresponding to the single-frame hyperspectral sample data of the changing time phase;

[0038] transforming the size of the variable phase hyperspectral single frame sample data according to the multiple different scale factors respectively, to obtain multiple scale-transformed variable phase hyperspectral single frame sample data;

[0039] cropping the image regions within the image cropping windows from the plurality of scale-transformed time-phase hyperspectral single sample data to obtain the plurality of images to be processed;

[0040] The method of cutting out a plurality of different image regions from the same single-frame hyperspectral sample data of the changing time phase to obtain a plurality of images to be processed further includes:

[0041] Using a plurality of differentiated spatial smoothing processing model parameters to perform spatial smoothing processing on the variable phase hyperspectral single sample data respectively, to obtain a plurality of processed variable phase hyperspectral single sample data;

[0042] Image regions within image cropping windows are cropped from the plurality of processed phase-varying hyperspectral single-frame sample data to obtain the plurality of images to be processed.

[0043] In one possible implementation, the graph attention feature encoder includes: a graph node feature embedding component and a graph topology attention component;

[0044] The extracting the remote sensing feature representation of the sample instance by the graph attention feature encoder that has completed the initial training includes:

[0045] For the target image to be processed in the sample instance, partitioning the target image to be processed by using a superpixel segmentation method to construct graph nodes, and obtaining multiple graph nodes corresponding to the same target image to be processed;

[0046] For a target graph node among the multiple graph nodes, performing feature embedding extraction on the target graph node by the graph node feature embedding component to obtain a spatial feature vector of the target graph node; wherein the spatial feature vector of the target graph node is used to indicate an image area of ​​the target graph node;

[0047] Determining a temporal phase feature vector of the target graph node and a topological feature vector of the target graph node; wherein the temporal phase feature vector of the target graph node is used to indicate the temporal phase to which the target graph node belongs, and the topological feature vector of the target graph node is used to indicate the spatial topological relationship of the target graph node in the image to be processed to which it belongs;

[0048] Merging the spatial feature vector of the target graph node, the temporal feature vector of the target graph node, and the topological feature vector of the target graph node to obtain a node spatiotemporal feature vector of the target graph node; the graph nodes corresponding to the multiple images to be processed of the same sample instance respectively use the same temporal feature vector;

[0049] Loading the node spatiotemporal feature vector corresponding to the sample instance into the graph topology attention component; wherein the node spatiotemporal feature vector corresponding to the sample instance includes: the node spatiotemporal feature vectors of the graph nodes corresponding to the multiple to-be-processed images belonging to the sample instance;

[0050] The graph topology attention component is used to perform attention aggregation on the node spatiotemporal feature vector corresponding to the sample instance to obtain the remote sensing feature representation corresponding to the sample instance.

[0051] In a possible implementation, the training of the intermediate remote sensing small sample change detection model using small sample change detection data examples to obtain the target remote sensing small sample change detection model includes:

[0052] The reference time phase feature discriminator and the change type discriminator are sequentially determined as feature discriminators to be updated; wherein the discriminators other than the feature discriminator in the reference time phase feature discriminator and the change type discriminator are determined as reference discrimination sources;

[0053] Determining a sample instance of the feature discriminator in the small sample change detection data instance;

[0054] Outputting, through the intermediate remote sensing small sample change detection model, a pending detection result and a baseline detection result corresponding to the sample instance of the feature discriminator; wherein the pending detection result is a detection result generated by the feature discriminator, and the baseline detection result is a detection result generated by the baseline discrimination source;

[0055] Determining the true detection result corresponding to the sample instance of the feature discriminator by using a preset true value discriminator; wherein the preset true value discriminator is a discriminator that has been trained and is used for the same discrimination task as the reference discrimination source;

[0056] determining an error parameter of the feature discriminator according to the pending detection result, the benchmark detection result, and the true detection result;

[0057] According to the error parameters of the feature discriminator, the model parameters of the graph attention feature encoder and the feature discriminator are updated until the model convergence conditions are met, thereby obtaining the target remote sensing small sample change detection model.

[0058] In a possible implementation, determining the error parameter of the feature discriminator according to the pending detection result, the benchmark detection result, and the actual detection result includes:

[0059] Determining a detection error of the feature discriminator based on the pending detection result and a target value of the change type corresponding to the pending detection result; wherein the detection error is used to characterize the reliability of the pending detection result generated by the feature discriminator;

[0060] Determining a baseline error of the feature discriminator based on the baseline detection result and the true detection result; wherein the baseline error is used to characterize the stability between the baseline detection result and the true detection result;

[0061] An error parameter of the feature discriminator is determined based on the detection error and the reference error.

[0062] In a possible implementation, the training of the intermediate remote sensing small sample change detection model using small sample change detection data examples to obtain the target remote sensing small sample change detection model includes:

[0063] Determining a plurality of sample instances of the intermediate remote sensing small sample change detection model from the small sample change detection data instance according to the spatiotemporal coverage distribution ratio of the baseline temporal phase remote sensing data instance and the change temporal phase remote sensing data instance;

[0064] Outputting a first detection result and a second detection result corresponding to the sample instance through the intermediate remote sensing small sample change detection model; wherein the first detection result is a detection result generated by the reference time phase feature discriminator, and the second detection result is a detection result generated by the change type discriminator;

[0065] determining a first detection error based on the first detection result and a target value of the change type corresponding to the first detection result, wherein the first detection error is used to characterize the reliability of the first detection result generated by the reference time phase feature discriminator;

[0066] determining a second detection error based on the second detection result and a change type target value corresponding to the second detection result, wherein the second detection error is used to characterize the reliability of the second detection result generated by the change type discriminator;

[0067] determining a first regularization term according to the first detection result and a true detection result corresponding to the first detection result, wherein the first regularization term is used to characterize the stability between the first detection result and the true detection result corresponding to the first detection result;

[0068] determining a second regularization term according to the second detection result and a true detection result corresponding to the second detection result, wherein the second regularization term is used to characterize the stability between the second detection result and the true detection result corresponding to the second detection result;

[0069] fusing the first detection error, the second detection error, the first regularization term, and the second regularization term according to a fusion coefficient to obtain a model error parameter; wherein the fusion coefficient is adaptively adjusted according to the error proportion of each discrimination task during training;

[0070] The model parameters of each network in the intermediate remote sensing small sample change detection model are updated according to the model error parameters until the model convergence conditions are met, thereby obtaining the target remote sensing small sample change detection model.

[0071] In a second aspect, an embodiment of the present invention provides a server system, including a server, wherein the server is configured to execute the method described in the first aspect.

[0072] Compared to existing technologies, the present invention provides the following beneficial effects: using a large-scale remote sensing small-sample change detection method and system incorporating graph attention, which relates to the field of artificial intelligence technology, the method includes: first, performing radiometric and atmospheric correction on the original remote sensing image to generate a standardized remote sensing image; aligning standardized images of different time phases to a unified geographic coordinate system through georeferencing; obtaining the aligned baseline and change phase images; and inputting the two-phase images into a pre-trained target remote sensing small-sample change detection model containing a graph attention feature encoder to output a change type map of the target area. The present invention captures the spatiotemporal topological relationships of ground objects through a graph attention mechanism, solving the problem of traditional methods relying on large amounts of labeled data and improving the accuracy of change detection in small-sample scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly describes the drawings required for use in the embodiments. It should be understood that the following drawings illustrate only certain embodiments of the present invention and should not be construed as limiting the scope of the present invention. Those skilled in the art can, without inventive effort, derive other relevant drawings from these drawings.

[0074] Figure 1 A schematic flow chart of the steps of a large-scale remote sensing small sample change detection method combined with graph attention provided by an embodiment of the present invention;

[0075] Figure 2 A schematic block diagram of the structure of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0076] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more apparent, the technical solutions of the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings of the embodiments of the present invention. It should be understood that the described embodiments are only a portion of the embodiments of the present invention, not all of them. Generally, the components of the embodiments of the present invention described and illustrated in the drawings herein may be arranged and designed in a variety of different configurations.

[0077] The specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0078] In order to solve the technical problems in the above background technology, Figure 1 This is a flow chart of a large-model remote sensing small sample change detection method combined with graph attention provided in an embodiment of the present disclosure. The large-model remote sensing small sample change detection method combined with graph attention is introduced in detail below.

[0079] Step S201, performing radiation correction and atmospheric correction on the original remote sensing image data of the target area to generate a standardized remote sensing image;

[0080] Step S202, aligning standardized remote sensing images of different time phases to a unified geographic coordinate system through georeferencing processing;

[0081] Step S203, obtaining aligned reference time phase remote sensing images and change time phase remote sensing images; the change time phase remote sensing images are remote sensing images acquired at a preset acquisition time interval of the reference time phase remote sensing images;

[0082] Step S204 : Inputting the reference time phase remote sensing image and the change time phase remote sensing image into a pre-trained target remote sensing small sample change detection model to obtain a change type map of the target area.

[0083] In an embodiment of the present invention, for example, the server conducts large-model remote sensing small sample change detection combined with graph attention on remote sensing image data of a target area (including multispectral images of baseline phase and change phase, with a spatial resolution of 10 meters, covering visible light to near-infrared bands).

[0084] First, the digital quantization (DN) values ​​of raw satellite imagery are affected by factors such as sensor gain offset, solar altitude, and atmospheric scattering / absorption, making them incapable of directly reflecting the true surface reflectance. The server first invokes the radiometric correction module of a specialized remote sensing processing tool, inputting the absolute calibration coefficients provided by the satellite (including gain and offset parameters for each band) to convert the DN values ​​into surface reflectance. For example, after processing, the DN value of a given pixel matches the true spectral characteristics of various features, such as vegetation (typically 30%-50%) and buildings (10%-20%). Subsequently, an atmospheric correction model (such as FLA ASH) is employed, combining the target area's atmospheric type (e.g., moderate continental atmosphere), visibility (e.g., 10 km, no severe haze), and an aerosol model (e.g., urban aerosol, suitable for anthropogenic emission scenarios such as industrial areas) to calculate atmospheric transmittance and aerosol scattering contributions, further eliminating atmospheric effects on the image. This ultimately generates a standardized remote sensing image (GeoTIFF file format), whose spectral information accurately reproduces surface conditions, laying the foundation for subsequent analysis.

[0085] Due to factors such as satellite orbital drift and topographical fluctuations, the standardized images of the baseline and change phases exhibit geometric differences (the position of the same feature in the two images may differ by approximately 2 pixels). Using the baseline phase image as a reference (which offers greater geometric accuracy), the server uses a feature extraction algorithm (such as SIFT) to extract 20 evenly distributed feature points with the same name, such as road intersections, building corners, and water boundary points (covering the four corners and center of the image to ensure uniform distribution). A polynomial model (such as a quadratic polynomial) is used to fit the coordinate relationships between the feature point pairs and construct a geometric transformation equation. Next, the change phase image is resampled using bilinear interpolation and converted to a coordinate system consistent with the baseline phase (such as the Transverse Mercator projection, suitable for mid-latitude regions). Finally, the registration accuracy is verified: the root mean square error (RMSE) for all feature point pairs is less than 1 pixel (approximately 10 meters), meeting the spatial consistency requirements for remote sensing change detection and ensuring that the spatial resolution and pixel arrangement of the two phases are fully consistent.

[0086] The server retrieves two standardized and georeferenced temporal images from the local remote sensing image database, confirming that they cover the same target area, use the same coordinate system, and have the same spatial resolution (10 meters) and pixel size (e.g., 512×512). At this point, the two temporal images meet the core requirements of "same area, same coordinate system, and same resolution" and can be directly used as input for subsequent change detection models.

[0087] The server concatenates the baseline phase (3 channels: red, green, and blue) and the change phase (3 channels) images into a 6-channel input tensor (with the same size as the original image, such as 512×512×6), normalizes the reflectance value to the range of 0-1 (for example, a reflectance of 0.2 corresponds to a normalized value of 0.2), and inputs the pre-trained "large model combined with graph attention" (this model is trained with small sample data, using only 10% of the labeled samples in the target area, solving the problem of traditional methods relying on large amounts of labeled data).

[0088] The internal processing logic of the model is as follows: First, each phase image is segmented into superpixels (e.g., using the SLIC algorithm, divided into 1000 superpixels), and the image is divided into multiple "graph nodes" (each node corresponds to a superpixel area with uniform area); then, the "spatial features" (the spectral average of all pixels in the area, such as the reflectivity average of the red, green, and blue bands), "phase features" (the baseline phase is marked as 0, the changing phase is marked as 1, distinguishing different time dimensions) and "topological features" of each node are extracted (the spatial Euclidean distance between the node and all other nodes, reflecting the spatial relationship between nodes), and the three are merged into a "node spatiotemporal feature vector" (with a dimension of about 1000+); then, the attention weights between nodes are calculated through the graph topology attention component (using Leaky The attention mechanism of the ReLU activation function reflects the spatiotemporal correlation between nodes. For example, the weights between nodes on the same road are higher, and the weights between corresponding nodes in different phases are higher). The node features are aggregated according to the weights (obtaining a "context feature vector" for each node, integrating information from surrounding nodes). Subsequently, the baseline phase feature discriminator (fully connected layer) outputs the feature type of each baseline phase node (such as buildings, roads, vegetation, and water bodies), and the change type discriminator (fully connected layer that integrates the context features of the two phases) outputs the change type of each node (such as new buildings, road expansions, vegetation reduction, and water body disappearance). Finally, the node-level prediction results are mapped back to the pixel level (for example, if a superpixel area contains 100 pixels, the change types of all pixels are the prediction results of the node) to generate a "change type map" (with the same size as the input image, such as 512×512).

[0089] The change type map uses pseudo-color coding (e.g., red for "new buildings," blue for "vegetation loss," green for "road expansion," and yellow for "water body disappearance") to clearly illustrate surface changes in the target area over time. For example, newly built areas within the target area appear in red, while areas with reduced vegetation appear in blue. The server saves this map as a GeoTIFF file containing coordinate system information, providing accurate spatial data support for applications such as urban planning (e.g., monitoring industrial expansion) and environmental protection (e.g., analyzing vegetation cover changes).

[0090] Throughout the entire process, from the original image to the final change type map, the server ensures spectral authenticity through standardized processing, spatial consistency through geographic registration, and captures spatiotemporal correlations through graph attention models. This enables accurate remote sensing change detection under small sample conditions, effectively solving the problem of traditional methods relying on large amounts of labeled data, and providing an efficient and intelligent solution for remote sensing image analysis.

[0091] In an embodiment of the present invention, the target remote sensing small sample change detection model is trained in the following manner.

[0092] Obtaining an initial remote sensing small sample change detection model, the initial remote sensing small sample change detection model comprising: a baseline temporal phase feature extraction model and a change discrimination model, the baseline temporal phase feature extraction model and the change discrimination model sharing a same graph attention feature encoder, the baseline temporal phase feature extraction model comprising the graph attention feature encoder and a baseline temporal phase ground feature discriminator, and the change discrimination model comprising the graph attention feature encoder and a change type discriminator;

[0093] The benchmark temporal phase feature extraction model is trained using a benchmark temporal phase remote sensing data instance, and model parameters of the image attention feature encoder and the benchmark temporal phase ground object discriminator included in the benchmark temporal phase feature extraction model are updated to obtain an image attention feature encoder and a benchmark temporal phase ground object discriminator that have completed initial training; wherein the benchmark temporal phase remote sensing data instance includes benchmark temporal phase remote sensing image data;

[0094] Based on the change type discriminator and the graph attention feature encoder that has completed initial training, constructing an initialized change discrimination model;

[0095] The change type discriminator in the initialized change discrimination model is trained using a change phase remote sensing data instance, freezing the model parameters of the graph attention feature encoder that has completed the initial training, and updating the model parameters of the change type discriminator to obtain a change type discriminator that has completed the initial training; wherein the change phase remote sensing data instance includes change phase remote sensing image data;

[0096] Constructing a change discrimination model that has completed initial training based on the graph attention feature encoder that has completed initial training and the change type discriminator that has completed initial training;

[0097] The change discrimination model that has completed the initial training is trained using the change phase remote sensing data instance, and the model parameters of the graph attention feature encoder that has completed the initial training and the change type discriminator that has completed the initial training are updated to obtain the graph attention feature encoder that has completed the advanced training and the change type discriminator that has completed the advanced training;

[0098] Based on the graph attention feature encoder that has completed advanced training, the baseline temporal feature discriminator that has completed initial training, and the change type discriminator that has completed advanced training, an intermediate remote sensing small sample change detection model is constructed;

[0099] The intermediate remote sensing small sample change detection model is trained using a small sample change detection data instance to obtain a target remote sensing small sample change detection model; wherein the output of the target remote sensing small sample change detection model is a change type map, and the small sample change detection data instance includes at least one of the following: the baseline phase remote sensing image data, the change phase remote sensing image data.

[0100] In an embodiment of the present invention, exemplarily, the initial model is composed of a baseline time phase feature extraction model and a change discrimination model, and the two share a graph attention feature encoder (used to capture the spatial topological relationship and spatiotemporal correlation of objects in remote sensing images). Among them, the baseline time phase feature extraction model includes a graph attention feature encoder and a baseline time phase object discriminator (outputting the baseline time phase object type, such as buildings, roads, etc.); the change discrimination model includes the same graph attention feature encoder and a change type discriminator (outputting the change type between the two time phases, such as new buildings, reduced vegetation, etc.). At this time, all model parameters are randomly initialized and no training is performed. The server uses a baseline time phase remote sensing data instance (including a baseline time phase multispectral image and corresponding object type label) to train the baseline time phase feature extraction model. During training, the graph attention feature encoder performs superpixel segmentation on the baseline phase image, dividing the image into several graph nodes. It extracts spatial features (such as spectral mean), topological features (such as spatial distance from other nodes), and temporal features (such as the baseline phase label) for each node. The node features are aggregated through an attention mechanism to generate a contextual feature vector. The baseline phase feature discriminator receives this feature vector and outputs a feature type prediction. The server updates the parameters of the graph attention feature encoder and the baseline phase feature discriminator through backpropagation, enabling the model to learn the feature representation and classification capabilities of features in the baseline phase. This results in an initially trained graph attention feature encoder (capable of feature feature extraction) and an initially trained baseline phase feature discriminator (capable of feature type discrimination). The server combines the initially trained graph attention feature encoder (with fixed parameters, preserving feature feature extraction) with the untrained change type discriminator (with randomly initialized parameters) to construct an initialized change discrimination model. At this point, the change type discriminator has not yet learned the characteristic patterns of change types and exists only as part of the model architecture. The server uses change phase remote sensing data instances (including change phase hyperspectral imagery and corresponding change type labels) to train the change type discriminator in the initialized change discrimination model. During training, the parameters of the graph attention feature encoder are frozen (not updated) and are solely responsible for extracting node context features from the change phase imagery. The change type discriminator receives these features and outputs a change type prediction. The server updates only the parameters of the change type discriminator through backpropagation, allowing the model to focus on learning the characteristic patterns of change types under the change phase. Ultimately, the initially trained change type discriminator (capable of distinguishing change types) is obtained. The server combines the initially trained graph attention feature encoder (with fixed parameters) with the initially trained change type discriminator (after updating parameters) to construct the initially trained change discrimination model. At this point, the model has preliminary change type detection capabilities, but the parameters of the graph attention feature encoder have not yet been optimized for change phase data. The server again uses change phase remote sensing data instances to train the initially trained change discrimination model.Unlike the previous step, this training step does not freeze the parameters of the graph attention feature encoder. Instead, the parameters of the encoder and the change type discriminator are updated simultaneously. During training, the graph attention feature encoder adjusts its feature extraction strategy based on the change phase data to generate contextual features more suitable for change type discrimination. The change type discriminator then optimizes change type predictions based on these updated features. Through the coordinated optimization of these two approaches, the final result is a graph attention feature encoder that has completed advanced training (with more accurate change-related feature extraction capabilities) and a change type discriminator that has completed advanced training (with more accurate change type discrimination capabilities). The server combines the graph attention feature encoder that has completed advanced training (with optimized feature extraction capabilities), the baseline phase feature discriminator that has completed initial training (with retained feature type discrimination capabilities), and the change type discriminator that has completed advanced training (with optimized change type discrimination capabilities) to construct an intermediate remote sensing small-sample change detection model. At this point, the model has integrated baseline phase feature discrimination and change type discrimination capabilities, but has not yet been tuned for small-sample change detection scenarios. The server trains an intermediate remote sensing small-sample change detection model using small-sample change detection data instances (consisting of a small number of base-phase and change-phase image pairs and corresponding change type labels). During training, the server fine-tunes model parameters (including the graph attention feature encoder, the base-phase feature discriminator, and the change type discriminator) to adapt the model to the small-sample change detection task. The graph attention feature encoder further optimizes feature extraction to capture subtle changes in small-sample data. The base-phase feature discriminator and the change type discriminator are collaboratively tuned to improve the accuracy of change type prediction. Ultimately, when the model converges, the target remote sensing small-sample change detection model is obtained. Its output is a change type map (pixel-level change type annotations), capable of accurately detecting surface changes under small-sample conditions. The entire training process is based on the principle of "feature extraction - discriminative ability learning - collaborative optimization - small-sample adaptation." By gradually iterating the model structure and parameters, the server fully utilizes the base-phase data (learning feature features), the change-phase data (learning change types), and the small-sample data (adapting to the target task), ultimately achieving a target model with small-sample change detection capabilities. The model captures the spatiotemporal correlation of objects through the graph attention mechanism, solving the problem of traditional methods relying on large amounts of labeled data, and can provide an efficient and intelligent solution for remote sensing image analysis.

[0101] In an embodiment of the present invention, the reference time phase remote sensing data instance is multispectral sequence data, and the reference time phase remote sensing data instance includes at least one reference time phase multispectral sequence sample data; the change time phase remote sensing data instance is change single frame data, and the change time phase remote sensing data instance includes at least one change time phase hyperspectral single frame sample data;

[0102] The method of using the benchmark time phase remote sensing data instance to train the benchmark time phase feature extraction model, updating the model parameters of the image attention feature encoder and the benchmark time phase land feature discriminator contained in the benchmark time phase feature extraction model, and obtaining the image attention feature encoder that has completed the initial training and the benchmark time phase land feature discriminator that has completed the initial training can be implemented through the following examples.

[0103] For each reference phase multispectral sequence sample data, a plurality of sequence time node images corresponding to the reference phase multispectral sequence sample data are obtained according to a plurality of extracted time sequence nodes;

[0104] Using a plurality of sequence time node images corresponding to the reference time phase multispectral sequence sample data, the reference time phase feature extraction model is trained, and the model parameters of the image attention feature encoder and the reference time phase ground object discriminator included in the reference time phase feature extraction model are updated to obtain the image attention feature encoder and the reference time phase ground object discriminator that have completed initial training;

[0105] The method of using the change phase remote sensing data instance to train the change type discriminator in the initialized change discrimination model, freezing the model parameters of the graph attention feature encoder that has completed the initial training, and updating the model parameters of the change type discriminator to obtain the change type discriminator that has completed the initial training, includes:

[0106] For each hyperspectral single-frame sample data of a changing time phase, data enhancement is performed on the hyperspectral single-frame sample data of the changing time phase to obtain a plurality of enhanced sample data corresponding to the hyperspectral single-frame sample data of the changing time phase;

[0107] The change type discriminator is trained using multiple enhanced sample data corresponding to the single-frame hyperspectral sample data of the change phase, the model parameters of the graph attention feature encoder that has completed the initial training are frozen, and the model parameters of the change type discriminator are updated to obtain the change type discriminator that has completed the initial training.

[0108] In an embodiment of the present invention, for example, the benchmark time phase remote sensing data instance read by the server is 1000 benchmark time phase multispectral sequence samples (each sample contains 4 multispectral images of the target area in spring (March), summer (June), autumn (September), and winter (December) of 2022, each image has 3 channels (red, green, and blue), and a spatial resolution of 10 meters), and each sample is annotated with a pixel-level ground object type label (four categories: buildings, roads, vegetation, and water bodies). For each sequence sample, the server directly extracts the 4 sequence time node images corresponding to the sample (i.e., images of March, June, September, and December 2022) according to the quarterly time series nodes (spring, summer, autumn, and winter), without the need for additional interpolation or screening. These node images cover the typical changes in the spectral characteristics of ground objects throughout the year (such as vegetation turning green in spring and falling leaves in autumn). During training, the server concatenates the four node images of each sequence sample into a 12-channel input tensor (4 phases × 3 bands) and feeds this into the baseline temporal feature extraction model. The graph attention feature encoder performs superpixel segmentation on each node image (SLIC algorithm, 1000 superpixels / image), extracts each node's spatial features (the mean pixel spectrum within the region), temporal features (phase label, such as 0 for spring and 1 for summer), and topological features (Euclidean distance from other nodes), and merges these features into a 1004-dimensional node spatiotemporal feature vector. A two-layer GAT (graph attention layer) calculates the attention weights between nodes (for example, nodes in the same vegetation area have higher weights), and aggregates them into a 128-dimensional contextual feature vector (incorporating spatiotemporal correlations). The baseline temporal feature discriminator (a two-layer fully connected network) receives the contextual feature vector and outputs the predicted probability of the feature type for each node (four categories). The server uses a cross-entropy loss function (calculating the error between the predicted probability and the true label) and backpropagates using the Adam optimizer (learning rate 1e-4, batch size 32). The server also updates the parameters of the graph attention feature encoder and the baseline temporal feature discriminator. Training lasts for 50 epochs, with the object classification accuracy on the validation set (200 samples) increasing from an initial 52% to 85% (loss value decreasing from 2.1 to 0.4). The server then saves the initially trained graph attention feature encoder (capable of extracting annual time series features) and the baseline temporal feature discriminator (with 85% accuracy). The server reads 500 single hyperspectral images of the target area in January 2023, covering 13 bands (visible to near-infrared) with a spatial resolution of 10 meters. Each image is annotated with a pixel-level change type label (new building, road expansion, vegetation loss, and water body disappearance).For each single sample, the server performs data augmentation to expand the small sample size: sliding window cropping: a 256×256 pixel window is slid across the sample with a 50% overlap, generating five subregion images for each sample (covering different spatial locations of the sample); scaling: each subregion image is scaled by 0.8x (reduction), 1.0x (original size), and 1.2x (enlargement), generating 3x data (5×3 = 15 total); spatial smoothing: the scaled image is smoothed using Gaussian filters with different sigma values ​​(0.5, 1.0, and 1.5) to simulate varying degrees of atmospheric scattering, generating 3x data (15×3 = 45 augmented samples per original sample). During training, the server feeds each augmented sample into a frozen graph attention feature encoder (with fixed parameters to preserve the ability to extract ground features), extracting a 128-dimensional contextual feature vector. This feature vector is then fed into a change type discriminator (a two-layer fully connected network with an input dimension of 128) to output a predicted probability of the change type (four categories). The server uses the cross-entropy loss function (calculating the error between the prediction and the true label) and the Adam optimizer (learning rate 5e-5, batch size 16) to update only the parameters of the change type discriminator (the encoder parameters are frozen to avoid destroying the learned land feature features). Training lasted for 30 rounds, and the change type detection accuracy of the validation set (500 enhanced samples) increased from the initial 31% to 72% (the loss value dropped from 2.3 to 0.6). Finally, the initially trained change type discriminator (with the ability to discriminate change types) was saved. Throughout the process, the server retained the interannual spectral change information of the baseline phase data through time series node extraction, allowing the encoder to learn more robust land feature features; multi-strategy data enhancement solved the small sample size problem of single-frame data in the change phase, allowing the change type discriminator to focus on learning change features. The synergy between the two laid the foundation for subsequent advanced model training.

[0109] In an embodiment of the present invention, the instance of the time-varying remote sensing data is a time-varying single-frame data, and the instance of the time-varying remote sensing data includes at least one time-varying hyperspectral single-frame sample data;

[0110] The method of using the change phase remote sensing data instance to train the change type discriminator in the initialized change discrimination model, freezing the model parameters of the graph attention feature encoder that has completed the initial training, and updating the model parameters of the change type discriminator to obtain the change type discriminator that has completed the initial training can be implemented through the following examples.

[0111] Cutting a plurality of different image regions from the same single-frame hyperspectral sample data of the changing time phase to obtain a plurality of images to be processed; wherein there are partially overlapping image regions between two of the images to be processed from the same single-frame hyperspectral sample data of the changing time phase;

[0112] determining the plurality of to-be-processed images as sample instances of the initialized change discrimination model;

[0113] Extracting a remote sensing feature representation of the sample instance through the graph attention feature encoder that has completed initial training;

[0114] Determining a change type detection result of the sample instance according to the remote sensing feature representation of the sample instance by the change type discriminator;

[0115] According to the change type detection result and the change type target value of the sample instance, an error parameter is determined, and based on the error parameter, the model parameter of the change type discriminator is updated to obtain the change type discriminator that has completed the initial training.

[0116] In an embodiment of the present invention, exemplarily, the server expands the sample size by local overlapping cropping for single-frame hyperspectral sample data of changing phase (taking the 13-band hyperspectral image of a certain industrial area in January 2023 as an example, with a spatial resolution of 10 meters, and marked with four types of pixel-level change type labels of "new building, road expansion, vegetation reduction, and water body disappearance"), and trains the change type discriminator in the initialized change discrimination model (the parameters of the image attention feature encoder are frozen). The specific process is as follows: the server reads a single-frame hyperspectral sample of changing phase (the file name is "20230120_industrial_hs.tif", the size is 1024×1024 pixels, and the coverage area is approximately 10.24 kilometers × 10.24 kilometers), and sets the cropping window parameters: the window size is 256×256 pixels (corresponding to an actual area of ​​approximately 2.56 kilometers × 2.56 kilometers), and the sliding step size is 128 pixels (overlap rate 50%). The server invokes the GDAL library's "Warp" tool, starting with the upper left corner of the sample and sliding the window horizontally and vertically to crop five images to be processed (named "20230120_industrial_hs_crop1.tif" through "_crop_5.tif"). These images overlap by 128 pixels (for example, crop 1 and crop 2 overlap by 128 pixels horizontally, and crop 1 and crop 3 overlap by 128 pixels vertically), ensuring that the sample covers the entire spatial area of ​​the original sample while preserving the spatial continuity of the features. The server directly uses the five cropped images as sample instances for initializing the change discrimination model (no additional preprocessing is required; only the 13-band hyperspectral information of the original sample is retained). Each sample instance has a size of 256 × 256 × 13, which meets the spatial dimension requirements of the model input (consistent with the input during baseline temporal phase training). The server calls the graph attention feature encoder that has completed initial training (parameters are frozen, retaining the ability to extract ground feature features) and performs feature extraction on each sample instance: Superpixel segmentation: The SLIC algorithm is used to segment each image to be processed into 1000 superpixels (each superpixel corresponds to an area of ​​approximately 65,500 square meters), ensuring that the segmentation result is consistent with the number of nodes during baseline phase training; Node feature fusion: The spatial features (13-band hyperspectral reflectance average, dimension 13), phase features (changing phase is marked as 1, dimension 1) and topological features (Euclidean distance with other 999 nodes, dimension 999) of each superpixel graph node are extracted and merged into a 1013-dimensional node spatiotemporal feature vector; Attention aggregation: The attention weights between nodes are calculated through a two-layer GAT (graph attention layer, 128 hidden units per layer) (for example, the nodes in the newly added building area have higher weights than the surrounding road nodes), and the 128-dimensional context feature vector is aggregated (integrating spatial topology and spatiotemporal correlation).The server inputs the context feature vector (128 dimensions) of each sample instance into the initialized change type discriminator (2-layer fully connected network, input dimension 128, output dimension 4), and outputs the predicted probability of the change type of each superpixel graph node (for example, the predicted probability of a node is "new building: 0.85, road expansion: 0.10, vegetation reduction: 0.03, water body disappearance: 0.02"). The server compares the predicted probability of the change type of each sample instance with the true label (the pixel-level change type label in the original sample) and uses the cross-entropy loss function to calculate the error parameter (for example, the loss value of a sample instance is 0.5, reflecting the deviation between the prediction and the true label). Subsequently, the server uses the Adam optimizer (learning rate 5e-5, batch size 16) to only update the parameters of the change type discriminator (the parameters of the graph attention feature encoder are frozen to avoid destroying the learned ground feature features). During training, the server uses 5-fold cross-validation (dividing the 2,500 processed images generated from the 500 original samples into five groups of 500 images each, which rotate as validation sets). After each round of training, the validation set's change type detection accuracy is calculated (for example, the accuracy is 31% in the first round, increases to 55% in the 10th round, and reaches 72% in the 30th round). Training stops when the validation set accuracy improves by less than 0.5% over three consecutive rounds, and the server saves the initially trained change type discriminator (after parameter update, the change type detection accuracy is 72%). The server generates multiple processed images from a single hyperspectral sample of the same change phase through local overlapping cropping, expanding the small sample size. A frozen graph attention feature encoder extracts stable ground features, ensuring that the change type discriminator focuses on learning change features. Cross-entropy loss and Adam optimization are used to update the discriminator parameters, ultimately resulting in a model component capable of discriminating change types. This entire process solves the small sample size issue for single-image data of a change phase, laying the foundation for subsequent advanced model training.

[0117] In the embodiment of the present invention, the step of cropping a plurality of different image regions from the same single-frame hyperspectral sample data of the changing time phase to obtain a plurality of images to be processed may be implemented through the following example.

[0118] Determine the size of the image cropping window;

[0119] The image cropping window is slid to a plurality of different sampling points of the hyperspectral single-frame sample data of the changing time phase, and the image areas within the image cropping window are cropped respectively to obtain the plurality of images to be processed.

[0120] In an embodiment of the present invention, for example, the server generates multiple images to be processed by sliding window cropping for a single sample of hyperspectral data of a changing phase (taking the 13-band hyperspectral image "20230120_industrial_hs.tif" of a certain industrial area in January 2023 as an example, with a size of 1024×1024 pixels, a spatial resolution of 10 meters, and a coverage area of ​​approximately 10.24 kilometers × 10.24 kilometers). The specific process is as follows: The server determines the cropping window size to be 256×256 pixels based on the model input consistency requirements (which must match the superpixel segmentation and feature extraction size during baseline phase training). This size corresponds to an actual geographic area of ​​approximately 2.56 kilometers × 2.56 kilometers, which can retain sufficient spatial information of the ground objects (such as complete building blocks and road sections) and meet the model's restrictions on the input spatial dimension (avoiding overload of computing resources due to excessive size). The server sets the sliding step size to 128 pixels (i.e., the window moves 128 pixels horizontally / vertically at a time, with a 50% overlap rate) to ensure that the spatial continuity of the features between adjacent images to be processed is preserved (for example, the same road or building area will not be completely divided due to cropping). Sampling point calculation: Horizontally, from left to right, the x-coordinates of the upper left corner are 0, 128, 256, 384, 512, 640, and 768 (a total of 7 sampling points, covering a sample width of 1024 pixels); vertically, from top to bottom, the y-coordinates of the upper left corner are also 0, 128, 256, 384, 512, 640, and 768 (a total of 7 sampling points, covering a sample height of 1024 pixels). Window sliding and cropping: The server loops through all sampling points (7×7=49 in total), slides the cropping window to each sampling point position (for example, the sampling point (0,0) corresponds to the window range x=0-255, y=0-255; the sampling point (128,0) corresponds to x=128-383, y=0-255), calls the "RasterIO" function of the GDAL library to crop the image area within the window, and saves each cropped image as an image to be processed (the naming format is "20230120_industrial_hs_crop_xx_yy.tif", where xx is the x-coordinate of the upper left corner, and yy is the y-coordinate of the upper left corner). Ultimately, 49 images were cropped from the same single hyperspectral sample of the change phase. Each image was 256 × 256 × 13 (13-band hyperspectral), and there was a 128-pixel overlap between each image (for example, "crop_0_0.tif" and "crop128_0.tif" overlapped horizontally by 128 pixels). This cropping method not only expanded the small sample size (converting one original sample into 49 samples to be processed) but also preserved the spatial topological relationships of the features, providing continuous and complete spatial feature data for the subsequent training of the change type discriminator.

[0121] In the embodiment of the present invention, the step of cropping a plurality of different image regions from the same single-frame hyperspectral sample data of the changing time phase to obtain a plurality of images to be processed may be implemented through the following example.

[0122] Determining the size of the image cropping window and its spatial layout position in the hyperspectral single sample data of the changing time phase;

[0123] Determining a plurality of different scale factors corresponding to the single-frame hyperspectral sample data of the changing time phase within a scale transformation range corresponding to the single-frame hyperspectral sample data of the changing time phase;

[0124] transforming the size of the variable phase hyperspectral single frame sample data according to the multiple different scale factors respectively, to obtain multiple scale-transformed variable phase hyperspectral single frame sample data;

[0125] The image regions within the image cropping window are respectively cropped from the multiple scale-transformed time-varying hyperspectral single-frame sample data to obtain the multiple images to be processed.

[0126] In an embodiment of the present invention, for example, the server generates multiple images to be processed through "scaling + fixed-position cropping" for a single hyperspectral sample of a change phase in a certain industrial zone in January 2023 (file name "20230120_industrial_hs.tif", size 1024×1024 pixels, 13 bands, spatial resolution 10 meters, and labeled with change type labels such as "new building"). The process is as follows: Based on the spatial consistency requirements of the model input (which must match the superpixel segmentation size trained on the benchmark phase), the server sets the cropping window size to 256×256 pixels (corresponding to an actual geographic area of ​​approximately 2.56 km×2.56 km, which can fully preserve typical features such as building blocks and road sections within the industrial zone). At the same time, to cover the core area of ​​the sample (the main production area of ​​the industrial zone), the sample center is selected as the spatial layout location, with the upper left corner coordinates of (384,384) (corresponding to the window range: x = 384-639 pixels, y = 384-639 pixels). The server uses a rescaling factor of 0.8 (downscaling), 1.0 (original size), and 1.2 (upscaling) within a rescaling range of 0.8-1.2. The downscaling factor simulates low-resolution satellite imagery (such as Sentinel-2's 10-meter resolution), while the upscaling factor simulates high-resolution imagery (such as Worldview-3's 0.3-meter resolution). This aims to enhance the model's robustness to changes in features at different scales (for example, the difference between the characteristics of a small building addition and a large road expansion). The server uses the "resize" function in the OpenCV library to rescale the original samples using bilinear interpolation (to maintain spectral continuity): a factor of 0.8 scales the original 1024×1024 samples to 819×819 pixels (1024×0.8, rounded to the nearest integer); a factor of 1.0 retains the original sample size (1024×1024); and a factor of 1.2 scales the original sample to 1229×1229 pixels (1024×1.2, rounded to the nearest integer). From each scaled sample, a 256×256 window is cropped according to the preset central spatial layout: 0.8 scaled sample (819×819): the center coordinates are (409,409) (819 / 2≈409), and the cropping window is 281-536×281-536 pixels (409±128, ensuring the window size is 256×256); 1.0 original sample (1024×1024): directly crop the preset central area (384-639×384-639 pixels); 1.2 enlarged sample (1229×1229): the center coordinates are (614,614) (1229 / 2≈614), and the cropping window is 486-741×486-741 pixels (614±128).Ultimately, the server generated three images to be processed (corresponding to scales 0.8, 1.0, and 1.2) from the same change phase sample. Each image had a 256×256×13 band size (preserving hyperspectral information) and covered the core industrial area of ​​the sample. This approach, through scaling, expanded the scale diversity of the sample, helping the change type discriminator learn change characteristics at different resolutions (such as the details of small-scale building additions in a zoomed-in image, or the overall trend of large-scale road expansion in a zoomed-out image), improving the model's generalization ability to real-world scenarios.

[0127] In the embodiment of the present invention, the step of cropping a plurality of different image regions from the same single-frame hyperspectral sample data of the changing time phase to obtain a plurality of images to be processed may be implemented through the following example.

[0128] Using a plurality of differentiated spatial smoothing processing model parameters to perform spatial smoothing processing on the variable phase hyperspectral single sample data respectively, to obtain a plurality of processed variable phase hyperspectral single sample data;

[0129] Image regions within image cropping windows are cropped from the plurality of processed phase-varying hyperspectral single-frame sample data to obtain the plurality of images to be processed.

[0130] In an embodiment of the present invention, for example, the server generates multiple images to be processed through "differential spatial smoothing + fixed window cropping" for a single sample of the changing phase hyperspectral image of an industrial area in January 2023 (the file name is "20230120_industrial_hs.tif", with a size of 1024×1024 pixels, 13 bands, a spatial resolution of 10 meters, and annotated with pixel-level change type labels such as "new building"). The process is as follows: the server selects a Gaussian filter (a commonly used spatial smoothing tool) and sets three differentiated sigma values ​​(standard deviation): 0.5, 1.0, and 1.5. The smaller the sigma value, the lighter the smoothing (retaining more features, such as building edges); the larger the sigma value, the heavier the smoothing (blurring details and highlighting the overall characteristics of large-area features, such as vegetation areas). These parameters are intended to simulate different degrees of atmospheric scattering or image noise to enhance the model's robustness to blurred images. The server calls the "GaussianBlur" function from the OpenCV library to batch process the original samples: sigma = 0.5: light smoothing is applied to each band, preserving details such as building edges and road markings (for example, the corners of newly added factories in industrial areas remain clearly visible); sigma = 1.0: moderate smoothing, balancing detail and noise (for example, the boundaries of road expansions retain continuity while eliminating local pixel noise); and sigma = 1.5: heavy smoothing, blurring small areas of detail (for example, small patches in areas of vegetation loss are merged into large areas, highlighting the overall trend). After processing, three processed hyperspectral single images (named "20230120_industrial_hs_sigma_0.5.tif," "_sigma1.0.tif," and "_sigma1.5.tif") are obtained, all retaining 13-band hyperspectral information. The server crops each processed sample using a preset fixed window (256×256 pixels, centered on the original sample's core industrial area, with a top-left corner at 384×384 pixels, corresponding to a window range of 384-639×384-639 pixels). For example, from the "sigma_0.5" sample, "20230120_industrial_hs_sigma_0.5_crop.tif" is cropped (preserving building edge details); from the "sigma1.5" sample, "20230120_industrial_hs_sigma1.5_crop.tif" is cropped (highlighting the overall trend of vegetation loss). Ultimately, the server generates three processed images (corresponding to three sigma values) from the same time-phase sample, each with 256×256×13 bands (preserving hyperspectral features).This method expands the fuzziness diversity of samples through differential spatial smoothing, helping the change type discriminator learn the change characteristics under different blur levels (such as the addition of small-area buildings in clear images and the reduction of large-area vegetation in blurred images), and improving the model's adaptability to interference such as "atmospheric scattering" and "image noise" in real scenes.

[0131] In an embodiment of the present invention, the graph attention feature encoder includes: a graph node feature embedding component and a graph topology attention component;

[0132] The extraction of the remote sensing feature representation of the sample instance by the graph attention feature encoder that has completed the initial training can be implemented through the following example.

[0133] For the target image to be processed in the sample instance, partitioning the target image to be processed by using a superpixel segmentation method to construct graph nodes, and obtaining multiple graph nodes corresponding to the same target image to be processed;

[0134] For a target graph node among the multiple graph nodes, performing feature embedding extraction on the target graph node by the graph node feature embedding component to obtain a spatial feature vector of the target graph node; wherein the spatial feature vector of the target graph node is used to indicate an image area of ​​the target graph node;

[0135] Determining a temporal phase feature vector of the target graph node and a topological feature vector of the target graph node; wherein the temporal phase feature vector of the target graph node is used to indicate the temporal phase to which the target graph node belongs, and the topological feature vector of the target graph node is used to indicate the spatial topological relationship of the target graph node in the image to be processed to which it belongs;

[0136] Merging the spatial feature vector of the target graph node, the temporal feature vector of the target graph node, and the topological feature vector of the target graph node to obtain a node spatiotemporal feature vector of the target graph node;

[0137] Loading the node spatiotemporal feature vector corresponding to the sample instance into the graph topology attention component; wherein the node spatiotemporal feature vector corresponding to the sample instance includes: the node spatiotemporal feature vectors of the graph nodes corresponding to the multiple to-be-processed images belonging to the sample instance;

[0138] The graph topology attention component is used to perform attention aggregation on the node spatiotemporal feature vector corresponding to the sample instance to obtain the remote sensing feature representation corresponding to the sample instance.

[0139] In an embodiment of the present invention, illustratively, the server extracts remote sensing feature representations for a single hyperspectral sample with a changing time phase (taking "20230120_industrial_hs_crop_0_0.tif" of an industrial area in January 2023 as an example, the image to be processed is a 256×256 pixel, 13-band hyperspectral image from the cropped area in the upper left corner of the original sample) through a graph attention feature encoder (including a graph node feature embedding component and a graph topology attention component). The process is as follows: the server calls the SLIC superpixel segmentation algorithm (setting the compactness to 10 and the number of segmentations to 1000) on the target image to be processed (256×256×13), and divides the image into 1000 continuous pixel areas (i.e., graph nodes). Each graph node corresponds to an actual geographic area of ​​approximately 65,500 square meters (256×256 pixels / 1000 nodes × 10 meters × 10 meters), covering typical terrain features such as "new buildings," "existing roads," and "vegetation reduction." For example, the graph node numbered "Node_001" is located in the upper left corner of the image and covers an area approximately the size of a standard factory building (approximately 60,000 square meters), including pixels representing new buildings. For each target graph node (such as "Node_001"), the graph node feature embedding component calculates the mean of its 13-band hyperspectral reflectance: it traverses all pixels within the node (approximately 65 pixels, 256×256 / 1000≈65) and takes the average value for each band (such as the visible red band and the near-infrared band) to obtain a 13-dimensional spatial feature vector. For example, the red band reflectance of "Node_001" has an average value of 0.15 (typical reflectance of buildings), and the near-infrared band has an average value of 0.20. The final spatial feature vector is [0.15, 0.12, 0.10, ..., 0.20] (13 values), which accurately indicates the spectral characteristics of the image area of ​​the node. Temporal feature vector: Since the image to be processed comes from a changing phase (January 2023), the server fixes its temporal feature vector to a 1-dimensional vector [1] (the base phase is [0]) to distinguish features in different time dimensions. Topological feature vector: The server calculates the spatial Euclidean distance (based on the actual distance after pixel coordinates are converted to geographic coordinates) between the target graph node (such as "Node_001") and the other 999 graph nodes. For example, "Node_001" is located in the upper left corner of the image (geographic coordinates: 120°15'00" E, 30°35'00" N), and "Node_500" is located in the center of the image (120°17'30" E, 30°32'30" N). The Euclidean distance between the two is approximately 3.5 kilometers. Therefore, the value corresponding to the position of "Node_500" in the topological feature vector of "Node_001" is 3500 meters. Ultimately, the topological feature vector is 999-dimensional, fully indicating the spatial topological relationship of the target graph nodes in the image to be processed.The server sequentially concatenates the spatial feature vector (13 dimensions), temporal feature vector (1 dimension), and topological feature vector (999 dimensions) of the target graph node to obtain a 1013-dimensional node spatiotemporal feature vector. For example, the node spatiotemporal feature vector for "Node_001" is [0.15, 0.12, ..., 0.20, 1, 0, 500, 1000, ..., 3500] (the first 13 bits are spatial features, the 14th bit is 1, and the last 999 bits are topological distances), integrating the spectral, temporal, and spatial location information of the node. The server collects the graph node spatiotemporal feature vectors for all images to be processed in the sample instance (for example, if the sample instance contains 49 cropped images, each with 1000 nodes, there are a total of 49 × 1000 = 49,000 node spatiotemporal feature vectors) and loads them into the graph topology attention component (a two-layer GAT network with 128 hidden units per layer). These vectors cover the entire spatial area of ​​the sample instance (1024×1024 pixels of the original sample) and preserve the spatial continuity of the features. The graph topology attention component calculates the correlation weights between nodes through the attention mechanism (using the Leaky-ReLU activation function). For example, the weight between "Node_001" (new building) and the adjacent "Node_002" (new building) is high (about 0.8), and the weight between it and the distant "Node_999" (vegetation reduction) is low (about 0.1). Subsequently, the component weightedly aggregates the node features according to the weights to obtain a 128-dimensional context feature vector for each node (incorporating the spatiotemporal features of neighboring nodes). For example, the context feature vector of "Node_001" incorporates the building features of "Node_002" and the road features of "Node_003", more comprehensively reflecting the feature combination in the area. Finally, the server aggregates the context feature vectors of all nodes to obtain the remote sensing feature representation corresponding to the sample instance (a 49,000×128-dimensional matrix). This representation preserves the spatiotemporal details of each graph node (such as the location and spectrum of newly added buildings) while also capturing spatial correlations between features (such as the adjacency between buildings and roads) through attention aggregation, providing precise feature input for the subsequent change type discriminator. Throughout the entire process, the server converts the image into a graph structure through superpixel segmentation, fuses spectral, temporal, and spatial information through feature embedding, and captures correlations between features through attention aggregation. Ultimately, it extracts a robust remote sensing feature representation, laying the foundation for small-sample change detection.

[0140] In the embodiment of the present invention, the graph nodes corresponding to the multiple images to be processed of the same sample instance respectively use the same temporal feature vector.

[0141] In the embodiment of the present invention, for example, the server processes 49 cropped images to be processed (all from January 2023) of the changing phase sample instance of "20230120_industrial_hs", and the graph nodes of all cropped images use the same phase feature vector [1] (indicating the changing phase) to ensure the consistency of the time dimension of the same instance.

[0142] In an embodiment of the present invention, the use of small sample change detection data instances to train the intermediate remote sensing small sample change detection model to obtain a target remote sensing small sample change detection model can be implemented through the following examples.

[0143] The reference time phase feature discriminator and the change type discriminator are sequentially determined as feature discriminators to be updated; wherein the discriminators other than the feature discriminator in the reference time phase feature discriminator and the change type discriminator are determined as reference discrimination sources;

[0144] Determining a sample instance of the feature discriminator in the small sample change detection data instance;

[0145] Outputting, through the intermediate remote sensing small sample change detection model, a pending detection result and a baseline detection result corresponding to the sample instance of the feature discriminator; wherein the pending detection result is a detection result generated by the feature discriminator, and the baseline detection result is a detection result generated by the baseline discrimination source;

[0146] Determining the true detection result corresponding to the sample instance of the feature discriminator by using a preset true value discriminator; wherein the preset true value discriminator is a discriminator that has been trained and is used for the same discrimination task as the reference discrimination source;

[0147] determining an error parameter of the feature discriminator according to the pending detection result, the benchmark detection result, and the true detection result;

[0148] According to the error parameters of the feature discriminator, the model parameters of the graph attention feature encoder and the feature discriminator are updated until the model convergence conditions are met, thereby obtaining the target remote sensing small sample change detection model.

[0149] In an embodiment of the present invention, for example, the server trains an intermediate remote sensing small sample change detection model (including a graph attention feature encoder GA T that has completed advanced training, a baseline phase feature discriminator D1 that has completed initial training, and a change type discriminator D2 that has completed advanced training) for a small sample change detection data instance of a certain industrial area (100 image pairs, including a multispectral image of the base phase in January 2022, a hyperspectral image of the change phase in January 2023, and corresponding pixel-level labels of the object type / change type). The process is as follows: the server sets D1 as the feature discriminator to be updated and D2 as the baseline discriminant source (both are intermediate model components). The base phase image (January 2022, 3 channels, 256×256 pixels) is extracted from the 100 small samples as a sample instance of D1 (a total of 100). After the baseline image is input into the intermediate model, GAT extracts 128-dimensional contextual features through superpixel segmentation (1000 nodes), feature embedding (spatial spectrum + phase labeling + topological distance) and attention aggregation; D1 receives the feature and outputs the pending detection result (land feature type prediction, such as the probability of a pixel being "building" is 0.85). At the same time, the corresponding change phase image (January 2023, 13 channels) is input into the intermediate model, and D2 outputs the baseline detection result (change type prediction, such as the probability of a pixel being "newly added building" is 0.7). To obtain the true detection result, the server calls the pre-trained high-accuracy change type discriminator T2 (trained with 1000 change phase samples, with an accuracy of 90%), inputs the change phase image into T2, and obtains a true result with a probability of "newly added building" of 0.95 (close to the sample label). Subsequently, the error is calculated using a fusion loss function: the weighted sum of the pending detection error of D1 (cross entropy loss with the true label of the baseline phase, weight 0.6) and the baseline detection error of D2 (cross entropy loss with the true result of T2, weight 0.4) (for example, the total error of a certain sample is 0.35). Backpropagation is performed using the Adam optimizer (learning rate 1e-6) to update the parameters of GAT and D1 (D2 is frozen). After 10 rounds of training, the accuracy of the validation set ground feature type increased from 80% to 85%. The server switches D2 as the feature discriminator to be updated, and D1 is used as the baseline discriminant source. From 100 small samples, the change phase image (January 2023, 13 channels, 256×256 pixels) is extracted as the sample instance of D2 (a total of 100). The change image is fed into the intermediate model, GAT extracts features, and D2 outputs a pending detection result (change type prediction, e.g., the probability of a pixel being "newly added building" is 0.75). Simultaneously, the corresponding baseline phase image is fed into the intermediate model, and D1 outputs a baseline detection result (ground feature type prediction, e.g., the probability of a pixel being "building" is 0.8). The pre-trained, high-accuracy ground feature type discriminator T1 (trained with 1000 baseline phase samples, with an accuracy of 92%) is called and fed into T1. The baseline image is then fed into T1, yielding a true result with a probability of 0.93 for "building."The error is calculated using a fusion loss function: the weighted sum of D2's pending detection error (cross entropy loss with the true label at the time of change, weighted 0.7) and D1's baseline detection error (cross entropy loss with the true result at T1, weighted 0.3) is calculated (for example, the total error for a given sample is 0.3). Backpropagation is performed using the Adam optimizer (learning rate 1e-6) to update the parameters of GAT and D2 (with D1 frozen). After 10 epochs of training, the validation set change type accuracy improved from 82% to 87%. The server iteratively updates D1 and D2, calculating the validation set accuracy for the target task (change type detection) after each epoch. Training is terminated when the accuracy improves by less than 0.5% for three consecutive epochs (e.g., 87% in epoch 20, 87.2% in epoch 21, and 87.3% in epoch 22), and the target remote sensing small-sample change detection model is saved. This model achieves 87% change type detection accuracy and accurately outputs change type maps for industrial areas (e.g., "new buildings" areas are marked in red, "vegetation loss" areas are marked in blue). The entire process alternately updates the two discriminators, using the output of the baseline discriminant source and the actual results of the preset true value discriminator to constrain the training of the discriminator to be updated, effectively improving the model performance in small sample scenarios and solving the problem of traditional methods relying on large amounts of labeled data.

[0150] In the embodiment of the present invention, determining the error parameter of the feature discriminator based on the pending detection result, the benchmark detection result and the actual detection result can be implemented through the following example.

[0151] Determining a detection error of the feature discriminator based on the pending detection result and a target value of the change type corresponding to the pending detection result; wherein the detection error is used to characterize the reliability of the pending detection result generated by the feature discriminator;

[0152] Determining a baseline error of the feature discriminator based on the baseline detection result and the true detection result; wherein the baseline error is used to characterize the stability between the baseline detection result and the true detection result;

[0153] An error parameter of the feature discriminator is determined based on the detection error and the reference error.

[0154] In an embodiment of the present invention, for example, the server calculates the error parameters of the intermediate model (the feature discriminator to be updated is the change type discriminator D2, the benchmark discriminant source is the benchmark phase feature discriminator D1, and the preset true value discriminator is the pre-trained high-accuracy feature discriminator T1) for a small sample change detection data instance (taking the base phase image of a certain industrial zone in January 2022, the change phase image in January 2023 and the corresponding "new building" pixel-level label sample as an example), and the process is as follows: the server extracts the error parameters from the sample The change phase image (January 2023, 13 bands, 256×256 pixels) is input into the graph attention feature encoder GAT of the intermediate model (advanced training is completed, and the parameters can be updated) to extract the 128-dimensional context feature vector; the feature discriminator D2 (change type discriminator) to be updated receives the feature and outputs the pending detection result. The predicted probability of the change type of a certain pixel is [0.8 (new building), 0.1 (road expansion), 0.05 (vegetation reduction), 0.05 (water body disappearance)]. The change type target value of the sample is the true label [1 (new building), 0, 0, 0] (the pixel is labeled "new building"). The server uses the cross entropy loss function to calculate the detection error (formula: Where (y i ) is the target value, (p i ) is the predicted probability): [L det = -[1×ln(0.8)+0×ln(0.1)+0×ln(0.05)+0×ln(0.05)]=-ln(0.8)≈0.223]; the detection error (0.223) represents the deviation between the pending detection result generated by D2 (probability of "new building" 0.8) and the true label. The smaller the value, the higher the reliability. The server extracts the baseline phase image (January 2022, 3 bands, 256×256 pixels) from the sample, inputs it into the GAT of the intermediate model, and extracts contextual features. The baseline discriminant source D1 (baseline phase feature discriminator, initial training completed, parameters frozen) receives this feature and outputs the baseline detection result. The predicted probability of the feature type of a certain pixel is [0.9 (building), 0.05 (road), 0.03 (vegetation), 0.02 (water body)] (the pixel is "building" in the baseline phase). To obtain the true detection results, the server calls the preset true value discriminator T1 (trained with 1000 benchmark phase samples, with a ground object classification accuracy of 92%), inputs the benchmark phase image into T1, and outputs the true detection results. The predicted probability of the ground object type of a certain pixel is [0.95 (building), 0.03 (road), 0.01 (vegetation), 0.01 (water body)] (close to the true label of the benchmark phase of the sample [1, 0, 0, 0]). The server uses the cross entropy loss function to calculate the benchmark error (the formula is the same as above, measuring the deviation between the benchmark detection result of D1 and the true detection result of T1): [L ref=-[0.95×ln(0.9)+0.03×ln(0.05)+0.01×ln(0.03)+0.01×ln(0.02)]≈-[0.95×(-0.105)+0.03×(-2.996)+0.01×(-3.507)+0.01×(-3.912)]≈0.051]; the benchmark error (0.051) represents the stability of the benchmark detection result of D1 (probability of "building" 0.9) and the true detection result of T1 (probability of "building" 0.95). The smaller the value, the more stable the output of D1. The server sets the detection error weight to 0.7 (reflecting the prediction reliability of the discriminator D2 to be updated) and the benchmark error weight to 0.3 (reflecting the output stability of the benchmark discriminant source D1) according to the task priority (change type detection is the target task and has a higher weight). The error parameter is obtained by weighting and fusing the two: [L total =0.7×L det t+0.3×L ref =0.7×0.223+0.3×0.051≈0.156+0.015=0.171]; This error parameter (0.171) combines the prediction accuracy of D2 (detection error) and the output stability of D1 (baseline error), ensuring that D2 can accurately predict the change type while constraining it to maintain consistency with the trained D1 (avoiding a decrease in feature discrimination ability due to updating D2). The server measures the prediction reliability of the discriminator to be updated using the detection error and the output stability of the benchmark discriminant source using the base error, then weightedly fuses the two to obtain the error parameter. This process ensures that when updating the feature discriminator, the performance of the target task (change type detection) is optimized while maintaining the overall stability of the model (without destroying the existing feature discrimination ability), effectively improving the model's generalization ability in small sample scenarios. For example, the error parameter of 0.171 of the above sample will be used for backpropagation to update the parameters of GAT and D2, so that the prediction probability of "new building" of D2 is increased from 0.8 to 0.9, which is closer to the true label (such as the error parameter drops to 0.1 after training).

[0155] In an embodiment of the present invention, the use of small sample change detection data instances to train the intermediate remote sensing small sample change detection model to obtain a target remote sensing small sample change detection model can be implemented through the following examples.

[0156] Determining a plurality of sample instances of the intermediate remote sensing small sample change detection model from the small sample change detection data instance according to the spatiotemporal coverage distribution ratio of the baseline temporal phase remote sensing data instance and the change temporal phase remote sensing data instance;

[0157] Outputting a first detection result and a second detection result corresponding to the sample instance through the intermediate remote sensing small sample change detection model; wherein the first detection result is a detection result generated by the reference time phase feature discriminator, and the second detection result is a detection result generated by the change type discriminator;

[0158] determining a model error parameter according to the first detection result and the second detection result;

[0159] The model parameters of each network in the intermediate remote sensing small sample change detection model are updated according to the model error parameters until the model convergence conditions are met, thereby obtaining the target remote sensing small sample change detection model.

[0160] In an embodiment of the present invention, for example, the server trains an intermediate remote sensing small sample change detection model (including a graph attention feature encoder GAT that has completed advanced training, a baseline phase ground feature discriminator D1 that has completed initial training, and a ground feature discriminator D2 that has completed advanced training) for a small sample change detection data instance of a certain area (including 100 "base phase + change phase" image pairs, the base phase is the multispectral images of the spring, summer, autumn and winter seasons in 2022, and the change phase is the hyperspectral images of the same period in 2023, all marked with pixel-level labels of ground feature type / change type) according to the following process. The trained change type discriminator D2) obtains the target model: the server first reviews the spatiotemporal coverage distribution of the baseline phase remote sensing data instances (four seasons in 2022, each accounting for 25%) and the change phase remote sensing data instances (the same period in 2023, each accounting for 25%): in terms of time, the baseline phase covers spring (March), summer (June), autumn (September), and winter (December), and the change phase is the same period in 2023; in terms of space, both the baseline and change phases cover industrial areas (40%), residential areas (30%), farmland (20%), and water bodies (10%). To maintain consistent training data distribution, the server selected four sets of sample instances (25 image pairs per set) from the 100 small samples: Spring: March 2022 baseline images + March 2023 change images (25 images, spatially covering 10 industrial areas, 7 residential areas, 5 farmland, and 3 water bodies); Summer: June 2022 baseline images + June 2023 change images (25 images, with the same spatial distribution); Autumn: September 2022 baseline images + September 2023 change images (25 images, with the same spatial distribution); Winter: December 2022 baseline images + December 2023 change images (25 images, with the same spatial distribution). The spatiotemporal distribution of each set of sample instances is identical to the previous training data, ensuring that the model can still learn comprehensive spatiotemporal features in small sample scenarios. The server inputs each group of sample instances into the intermediate model and processes the baseline phase image and the change phase image in turn: Baseline phase image processing: Taking the "Industrial Zone Baseline Image in March 2022" (3 channels, 256×256 pixels) as an example, GAT extracts the spatial features (spectral mean), phase features (spring is marked as 0), and topological features (Euclidean distance with other nodes) of each node through superpixel segmentation (1000 nodes), and merges them into a 1013-dimensional node spatiotemporal feature vector; 128-dimensional contextual features are obtained through attention aggregation (2-layer GAT); the baseline phase feature discriminator D1 (2-layer fully connected network) receives the feature and outputs the first detection result. The predicted probability of the feature type of a certain pixel is [0.85 (building), 0.08 (road), 0.05 (vegetation), 0.02 (water body)] (the true label of the pixel in the baseline phase is "building").Processing of changing phase images: Taking the "Industrial Zone Change Image in March 2023" (13 channels, 256×256 pixels) as an example, GAT extracts features of the same structure (the phase feature is marked as 1 for spring); the change type discriminator D2 (2-layer fully connected network) receives the features and outputs the second detection result. The predicted probability of the change type of a certain pixel is [0.75 (new building), 0.15 (road expansion), 0.07 (vegetation reduction), 0.03 (disappearance of water body)] (the true label of this pixel is "new building"). For each sample instance, the server fuses the first detection result error (D1's ground feature discrimination error) and the second detection result error (D2's change type discrimination error) to obtain the model error parameter: First detection result error: The cross entropy loss function is used to calculate the deviation between the predicted probability of D1 and the true label of the baseline phase (Formula: ). y 1i is the ground truth label, p 1i is the predicted probability of D1). For example, the true label of the above-mentioned reference image pixels is [1 (building), 0, 0, 0], then: [L1 = -[1×ln(0.85)+0×ln(0.08)+0×ln(0.05)+0×ln(0.02)] = -ln(0.85)≈0.17]; The second detection result error: The cross entropy loss function is also used to calculate the deviation between the predicted probability of D2 and the true label of the change type (Formula: y 2i is the true label of the change type, p 2i is the predicted probability of D2). For example, if the true label of the above-mentioned change image pixels is [1 (new building), 0, 0, 0], then: [L2 = -[1×ln(0.75)+0×ln(0.15)+0×ln(0.07)+0×ln(0.03)] = -ln(0.75)≈0.29]; Model error parameters: According to the priority of the target task (change type detection is the core task), the first error weight is set to 0.3 (ground feature discrimination is auxiliary), the second error weight is 0.7 (change type is the target), and the weighted sum fusion is adopted: [L total =0.3×L1+0.7×L2=0.3×0.17+0.7×0.29=0.051+0.203=0.254]; the server uses the Adam optimizer (learning rate 1e-6, batch size 8) to reduce the model error parameter L total Back propagation is performed to update the parameters of GAT, D1, and D2 simultaneously (all components of the intermediate model are fine-tuned to ensure the coordinated optimization of feature extraction and discrimination capabilities). After a round of training, the server uses a validation set (20 samples, with the same temporal and spatial distribution as the training set) to evaluate the accuracy of change type detection (target task performance indicator). For example: After the first round of training, the validation set accuracy is 82% (L totalThe average is 0.32); after the fifth round of training, the accuracy rate increased to 85% (L total The average is 0.21); after the 10th round of training, the accuracy is 85.3% (L total The average is 0.19); after the 11th round of training, the accuracy is 85.4% (L total The average is 0.188); after the 12th round of training, the accuracy is 85.5% (L total The average is 0.185). When the validation set accuracy improves by less than 0.5% for three consecutive rounds (0.2% in rounds 10-12), the server determines that the model has converged, stops training, and saves the target remote sensing small sample change detection model. The model's change type detection accuracy reaches 85.5%, and it can accurately output regional change type maps (such as "new buildings" in industrial areas are marked in red, and "vegetation reduction" areas in farmland are marked in blue). The server selects sample instances based on the temporal and spatial coverage distribution ratio to ensure that the small sample data is consistent with the previous training data distribution; outputs two detection results (ground objects and change types) through the intermediate model, and fuses the errors of the two to obtain the model error parameters; updates all component parameters through backpropagation until the target task performance converges. The entire process effectively solves the "data distribution offset" problem in small sample scenarios, improves the generalization ability of the model, and ultimately obtains the target model that can accurately detect surface changes.

[0161] In the embodiment of the present invention, determining the model error parameter according to the first detection result and the second detection result can be implemented through the following example.

[0162] determining a first detection error based on the first detection result and a target value of the change type corresponding to the first detection result, wherein the first detection error is used to characterize the reliability of the first detection result generated by the reference time phase feature discriminator;

[0163] determining a second detection error based on the second detection result and a change type target value corresponding to the second detection result, wherein the second detection error is used to characterize the reliability of the second detection result generated by the change type discriminator;

[0164] determining a first regularization term according to the first detection result and a true detection result corresponding to the first detection result, wherein the first regularization term is used to characterize the stability between the first detection result and the true detection result corresponding to the first detection result;

[0165] determining a second regularization term according to the second detection result and a true detection result corresponding to the second detection result, wherein the second regularization term is used to characterize the stability between the second detection result and the true detection result corresponding to the second detection result;

[0166] According to the fusion coefficient, the first detection error, the second detection error, the first regularization term and the second regularization term are fused to obtain the model error parameter; wherein the fusion coefficient is adaptively adjusted according to the error proportion of each discrimination task during training.

[0167] In an embodiment of the present invention, exemplarily, the server calculates the model error parameters of the intermediate remote sensing small sample change detection model (including the graph attention feature encoder GAT, the baseline phase feature discriminator D1, and the change type discriminator D2) for a small sample change detection data instance of an industrial area (including the multispectral image of the baseline phase in January 2022, the hyperspectral image of the change phase in January 2023, and the corresponding pixel-level labels of "building" (baseline feature) and "new building" (change type)). The process is as follows: the server inputs the baseline phase image (January 2022, 3 channels, 256×256 pixels) into the intermediate model, GAT extracts the node spatiotemporal features (spatial spectrum + phase label + topological distance) and aggregates them into context features, D1 outputs the first detection result, and the predicted probability of the feature type of a pixel is [0.85 (building), 0.08 (road), 0.05 (vegetation), 0.02 (water body)] (the true label of the pixel baseline phase is "building"). At the same time, the change phase image (January 2023, 13 channels, 256×256 pixels) is input into the intermediate model, GAT extracts features, and D2 outputs the second detection result. The predicted probability of the change type of a certain pixel is [0.75 (new building), 0.15 (road expansion), 0.07 (vegetation reduction), 0.03 (water body disappearance)] (the true label of the pixel change type is "new building"). First detection error (reliability of D1): The server uses the cross entropy loss function to calculate the deviation between the first detection result of D1 and the true label of the baseline phase (formula: y 1i is the ground truth label, p 1i is the predicted probability of D1). The true label of this pixel is [1(building), 0, 0, 0], so: L det1 =-[1×ln(0.85)+0×ln(0.08)+0×ln(0.05)+0×ln(0.02)]=-ln(0.85)≈0.17; This value represents the prediction reliability of D1 for the benchmark feature (the smaller the value, the higher the reliability). Second detection error (reliability of D2): Also using cross entropy loss, calculate the deviation between the second detection result of D2 and the true label of the change type (formula: y 2i is the true label of the change type, p 2i is the predicted probability of D2). The true label of this pixel is [1 (new building), 0, 0, 0], so: L det2=-[1×ln(0.75)+0×ln(0.15)+0×ln(0.07)+0×ln(0.03)]=-ln(0.75)≈0.29; This value represents the reliability of D2's prediction of the change type (the smaller the value, the higher the reliability). To ensure the stability of the model output (not deviating from the trained reliable results), the server calls the preset true value discriminator (pre-trained high-accuracy model): The first regularization term (stability of D1): Call the pre-trained ground feature discriminator T1 (trained with 1000 benchmark time phase samples, with an accuracy of 92%), input the benchmark time phase image into T1, and obtain the true detection result. The predicted probability of the ground feature type of a certain pixel is [0.95 (building), 0.03 (road), 0.01 (vegetation), 0.01 (water body)] (close to the true label). The cross entropy is used to calculate the deviation between the first detection result of D1 and the output of T1 (formula: t 1i is the true detection result output by T1): R1 = -[0.95×ln(0.85)+0.03×ln(0.08)+0.01×ln(0.05)+0.01×ln(0.02)]≈0.05; this value represents the stability of the output of D1 and the output of T1 (the smaller the value, the more stable D1). The second regularization term (stability of D2): call the pre-trained change type discriminator T2 (trained with 1000 change phase samples, with an accuracy of 90%), input the change phase image into T2, and obtain the true detection result. The predicted probability of the change type of a certain pixel is [0.9 (new building), 0.07 (road expansion), 0.02 (vegetation reduction), 0.01 (water body disappearance)] (close to the true label). The cross entropy is used to calculate the deviation between the second detection result of D2 and the output of T2 (formula: t 2i (where R2 is the true detection result output by T2): R2 = -[0.9×ln(0.75)+0.07×ln(0.15)+0.02×ln(0.07)+0.01×ln(0.03)]≈0.12; this value indicates the stability of the D2 output relative to the T2 output (the smaller the value, the more stable D2). The server adaptively adjusts the fusion coefficient based on the error ratio during training (ensuring a higher error ratio for the target task (change type detection)). For example, in the previous round of training, the second detection error (D2's change type error) accounted for 60% of the total error, the first detection error accounted for 30%, and the regularization term accounted for 10%. Therefore, the fusion coefficients for this round are set as follows: (α1 = 0.2) (first detection error, D1's reliability); (α2 = 0.5) (second detection error, D2's reliability, target task, highest weight); (α3 = 0.15) (first regularization term, D1's stability); (α4 = 0.15) (second regularization term, D2's stability). By weighting and fusing the four terms, we get the model error parameter: Ltotal =α1×L det1 +α2×L det2 +α3×R1+α4×R2=0.2×0.17+0.5×0.29+0.15×0.05+0.15×0.12=0.034+0.145+0.0075+0.018=0.2045; the server measures the discriminator's prediction reliability using detection error (D1 for the baseline feature, D2 for the change type), and uses the regularization term to measure the stability of the discriminator output and the pre-trained true value discriminator. The fusion coefficient is then adaptively adjusted based on the error proportion, and the four terms are weighted and fused to obtain the model error parameter. This process not only ensures the performance optimization of the target task (change type detection) (the second detection error has the highest weight), but also constrains the stability of the model output (regularization term), effectively improving the model's generalization ability in small sample scenarios. For example, the model error parameter of 0.2045 of the above sample will be used for backpropagation to update the parameters of GAT, D1, and D2, so that the prediction probability of "new building" of D2 is increased from 0.75 to 0.9, which is closer to the output of T2 (such as the error parameter drops to 0.15 after training).

[0168] The embodiment of the present invention provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the aforementioned large-scale remote sensing small sample change detection method combined with graph attention. Figure 2 As shown, Figure 2 This is a block diagram of the structure of a computer device 100 provided in an embodiment of the present invention. Computer device 100 includes a memory 111, a processor 112, and a communication unit 113. To enable data transmission or exchange, memory 111, processor 112, and communication unit 113 are electrically connected to each other, directly or indirectly. For example, these components can be electrically connected via one or more communication buses or signal lines.

[0169] For illustrative purposes, the foregoing description has been made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the present disclosure to the precise forms disclosed. Numerous modifications and variations are possible in light of the above teachings. These embodiments have been selected and described in order to best illustrate the principles of the present disclosure and its practical application, thereby enabling those skilled in the art to best utilize the present disclosure and to utilize various embodiments with various modifications as appropriate for the specific application contemplated.

Claims

1. A large-scale remote sensing small-sample change detection method based on graph attention, characterized in that: include: Perform radiation correction and atmospheric correction on the original remote sensing image data of the target area to generate standardized remote sensing images; Through georeferencing, standardized remote sensing images of different phases are aligned to a unified geographic coordinate system; Acquire aligned reference time phase remote sensing images and change time phase remote sensing images; the change time phase remote sensing images are remote sensing images acquired at a preset acquisition time interval of the reference time phase remote sensing images; The reference time phase remote sensing image and the change time phase remote sensing image are input into a pre-trained target remote sensing small sample change detection model to obtain a change type map of the target area.

2. The method according to claim 1, characterized in that The target remote sensing small sample change detection model is trained by the following methods, including: Obtaining an initial remote sensing small sample change detection model, the initial remote sensing small sample change detection model comprising: a baseline temporal phase feature extraction model and a change discrimination model, the baseline temporal phase feature extraction model and the change discrimination model sharing a same graph attention feature encoder, the baseline temporal phase feature extraction model comprising the graph attention feature encoder and a baseline temporal phase ground feature discriminator, and the change discrimination model comprising the graph attention feature encoder and a change type discriminator; The benchmark temporal phase feature extraction model is trained using a benchmark temporal phase remote sensing data instance, and model parameters of the image attention feature encoder and the benchmark temporal phase ground object discriminator included in the benchmark temporal phase feature extraction model are updated to obtain an image attention feature encoder and a benchmark temporal phase ground object discriminator that have completed initial training; wherein the benchmark temporal phase remote sensing data instance includes benchmark temporal phase remote sensing image data; Based on the change type discriminator and the graph attention feature encoder that has completed initial training, constructing an initialized change discrimination model; The change type discriminator in the initialized change discrimination model is trained using a change phase remote sensing data instance, freezing the model parameters of the graph attention feature encoder that has completed the initial training, and updating the model parameters of the change type discriminator to obtain a change type discriminator that has completed the initial training; wherein the change phase remote sensing data instance includes change phase remote sensing image data; Constructing a change discrimination model that has completed initial training based on the graph attention feature encoder that has completed initial training and the change type discriminator that has completed initial training; The change discrimination model that has completed the initial training is trained using the change phase remote sensing data instance, and the model parameters of the graph attention feature encoder that has completed the initial training and the change type discriminator that has completed the initial training are updated to obtain the graph attention feature encoder that has completed the advanced training and the change type discriminator that has completed the advanced training; Based on the graph attention feature encoder that has completed advanced training, the baseline temporal feature discriminator that has completed initial training, and the change type discriminator that has completed advanced training, an intermediate remote sensing small sample change detection model is constructed; The intermediate remote sensing small sample change detection model is trained using a small sample change detection data instance to obtain a target remote sensing small sample change detection model; wherein the output of the target remote sensing small sample change detection model is a change type map, and the small sample change detection data instance includes at least one of the following: the baseline phase remote sensing image data, the change phase remote sensing image data.

3. The method according to claim 2, characterized in that The reference time phase remote sensing data instance is multispectral sequence data, and the reference time phase remote sensing data instance includes at least one reference time phase multispectral sequence sample data; the change time phase remote sensing data instance is change single frame data, and the change time phase remote sensing data instance includes at least one change time phase hyperspectral single frame sample data; The method of training the benchmark temporal phase feature extraction model using a benchmark temporal phase remote sensing data example, updating the model parameters of the graph attention feature encoder and the benchmark temporal phase ground object discriminator contained in the benchmark temporal phase feature extraction model, and obtaining the graph attention feature encoder and the benchmark temporal phase ground object discriminator that have completed initial training, includes: For each reference phase multispectral sequence sample data, a plurality of sequence time node images corresponding to the reference phase multispectral sequence sample data are obtained according to a plurality of extracted time sequence nodes; Using a plurality of sequence time node images corresponding to the reference time phase multispectral sequence sample data, the reference time phase feature extraction model is trained, and the model parameters of the image attention feature encoder and the reference time phase ground object discriminator included in the reference time phase feature extraction model are updated to obtain the image attention feature encoder and the reference time phase ground object discriminator that have completed initial training; The method of using the change phase remote sensing data instance to train the change type discriminator in the initialized change discrimination model, freezing the model parameters of the graph attention feature encoder that has completed the initial training, and updating the model parameters of the change type discriminator to obtain the change type discriminator that has completed the initial training, includes: For each hyperspectral single-frame sample data of a changing time phase, data enhancement is performed on the hyperspectral single-frame sample data of the changing time phase to obtain a plurality of enhanced sample data corresponding to the hyperspectral single-frame sample data of the changing time phase; The change type discriminator is trained using multiple enhanced sample data corresponding to the single-frame hyperspectral sample data of the change phase, the model parameters of the graph attention feature encoder that has completed the initial training are frozen, and the model parameters of the change type discriminator are updated to obtain the change type discriminator that has completed the initial training.

4. The method according to claim 2, characterized in that The instance of the time-varying remote sensing data is a single-frame data, and the instance of the time-varying remote sensing data includes at least one single-frame sample data of a hyperspectral change phase; The method of using the change phase remote sensing data instance to train the change type discriminator in the initialized change discrimination model, freezing the model parameters of the graph attention feature encoder that has completed the initial training, and updating the model parameters of the change type discriminator to obtain the change type discriminator that has completed the initial training, includes: Cutting a plurality of different image regions from the same single-frame hyperspectral sample data of the changing time phase to obtain a plurality of images to be processed; wherein there are partially overlapping image regions between two of the images to be processed from the same single-frame hyperspectral sample data of the changing time phase; determining the plurality of to-be-processed images as sample instances of the initialized change discrimination model; Extracting a remote sensing feature representation of the sample instance through the graph attention feature encoder that has completed initial training; Determining a change type detection result of the sample instance according to the remote sensing feature representation of the sample instance by the change type discriminator; According to the change type detection result and the change type target value of the sample instance, an error parameter is determined, and based on the error parameter, the model parameter of the change type discriminator is updated to obtain the change type discriminator that has completed the initial training.

5. The method according to claim 4, characterized in that The method of cutting out a plurality of different image regions from the same single-frame hyperspectral sample data of the changing time phase to obtain a plurality of images to be processed includes: Determine the size of the image cropping window; Sliding the image cropping window to a plurality of different sampling points of the hyperspectral single sample data of the changing time phase, cropping the image areas within the image cropping window respectively, to obtain the plurality of images to be processed; The method of cutting out a plurality of different image regions from the same single-frame hyperspectral sample data of the changing time phase to obtain a plurality of images to be processed further includes: Determining the size of the image cropping window and its spatial layout position in the hyperspectral single sample data of the changing time phase; Determining a plurality of different scale factors corresponding to the single-frame hyperspectral sample data of the changing time phase within a scale transformation range corresponding to the single-frame hyperspectral sample data of the changing time phase; transforming the size of the variable phase hyperspectral single frame sample data according to the multiple different scale factors respectively, to obtain multiple scale-transformed variable phase hyperspectral single frame sample data; cropping the image regions within the image cropping windows from the plurality of scale-transformed time-phase hyperspectral single sample data to obtain the plurality of images to be processed; The method of cutting out a plurality of different image regions from the same single-frame hyperspectral sample data of the changing time phase to obtain a plurality of images to be processed further includes: Using a plurality of differentiated spatial smoothing processing model parameters to perform spatial smoothing processing on the variable phase hyperspectral single sample data respectively, to obtain a plurality of processed variable phase hyperspectral single sample data; Image regions within image cropping windows are cropped from the plurality of processed phase-varying hyperspectral single-frame sample data to obtain the plurality of images to be processed.

6. The method according to claim 4, characterized in that The graph attention feature encoder includes: a graph node feature embedding component and a graph topology attention component; The extracting the remote sensing feature representation of the sample instance by the graph attention feature encoder that has completed the initial training includes: For the target image to be processed in the sample instance, partitioning the target image to be processed by using a superpixel segmentation method to construct graph nodes, and obtaining multiple graph nodes corresponding to the same target image to be processed; For a target graph node among the multiple graph nodes, performing feature embedding extraction on the target graph node by the graph node feature embedding component to obtain a spatial feature vector of the target graph node; wherein the spatial feature vector of the target graph node is used to indicate an image area of ​​the target graph node; Determining a temporal phase feature vector of the target graph node and a topological feature vector of the target graph node; wherein the temporal phase feature vector of the target graph node is used to indicate the temporal phase to which the target graph node belongs, and the topological feature vector of the target graph node is used to indicate the spatial topological relationship of the target graph node in the image to be processed to which it belongs; Merging the spatial feature vector of the target graph node, the temporal feature vector of the target graph node, and the topological feature vector of the target graph node to obtain a node spatiotemporal feature vector of the target graph node; the graph nodes corresponding to the multiple images to be processed of the same sample instance respectively use the same temporal feature vector; Loading the node spatiotemporal feature vector corresponding to the sample instance into the graph topology attention component; wherein the node spatiotemporal feature vector corresponding to the sample instance includes: the node spatiotemporal feature vectors of the graph nodes corresponding to the multiple to-be-processed images belonging to the sample instance; The graph topology attention component is used to perform attention aggregation on the node spatiotemporal feature vector corresponding to the sample instance to obtain the remote sensing feature representation corresponding to the sample instance.

7. The method according to claim 2, characterized in that The method of training the intermediate remote sensing small sample change detection model using the small sample change detection data instance to obtain the target remote sensing small sample change detection model includes: The reference time phase feature discriminator and the change type discriminator are sequentially determined as feature discriminators to be updated; wherein the discriminators other than the feature discriminator in the reference time phase feature discriminator and the change type discriminator are determined as reference discrimination sources; Determining a sample instance of the feature discriminator in the small sample change detection data instance; Outputting, through the intermediate remote sensing small sample change detection model, a pending detection result and a baseline detection result corresponding to the sample instance of the feature discriminator; wherein the pending detection result is a detection result generated by the feature discriminator, and the baseline detection result is a detection result generated by the baseline discrimination source; Determining the true detection result corresponding to the sample instance of the feature discriminator by using a preset true value discriminator; wherein the preset true value discriminator is a discriminator that has been trained and is used for the same discrimination task as the reference discrimination source; determining an error parameter of the feature discriminator according to the pending detection result, the benchmark detection result, and the true detection result; According to the error parameters of the feature discriminator, the model parameters of the graph attention feature encoder and the feature discriminator are updated until the model convergence conditions are met, thereby obtaining the target remote sensing small sample change detection model.

8. The method according to claim 7, characterized in that The determining of the error parameter of the feature discriminator according to the pending detection result, the benchmark detection result, and the true detection result includes: Determining a detection error of the feature discriminator based on the pending detection result and a target value of the change type corresponding to the pending detection result; wherein the detection error is used to characterize the reliability of the pending detection result generated by the feature discriminator; Determining a baseline error of the feature discriminator based on the baseline detection result and the true detection result; wherein the baseline error is used to characterize the stability between the baseline detection result and the true detection result; An error parameter of the feature discriminator is determined based on the detection error and the reference error.

9. The method according to claim 2, characterized in that The method of training the intermediate remote sensing small sample change detection model using the small sample change detection data instance to obtain the target remote sensing small sample change detection model includes: Determining a plurality of sample instances of the intermediate remote sensing small sample change detection model from the small sample change detection data instance according to the spatiotemporal coverage distribution ratio of the baseline temporal phase remote sensing data instance and the change temporal phase remote sensing data instance; Outputting a first detection result and a second detection result corresponding to the sample instance through the intermediate remote sensing small sample change detection model; wherein the first detection result is a detection result generated by the reference time phase feature discriminator, and the second detection result is a detection result generated by the change type discriminator; determining a first detection error based on the first detection result and a target value of the change type corresponding to the first detection result, wherein the first detection error is used to characterize the reliability of the first detection result generated by the reference time phase feature discriminator; determining a second detection error based on the second detection result and a change type target value corresponding to the second detection result, wherein the second detection error is used to characterize the reliability of the second detection result generated by the change type discriminator; determining a first regularization term according to the first detection result and a true detection result corresponding to the first detection result, wherein the first regularization term is used to characterize the stability between the first detection result and the true detection result corresponding to the first detection result; determining a second regularization term according to the second detection result and a true detection result corresponding to the second detection result, wherein the second regularization term is used to characterize the stability between the second detection result and the true detection result corresponding to the second detection result; fusing the first detection error, the second detection error, the first regularization term, and the second regularization term according to a fusion coefficient to obtain a model error parameter; wherein the fusion coefficient is adaptively adjusted according to the error proportion of each discrimination task during training; The model parameters of each network in the intermediate remote sensing small sample change detection model are updated according to the model error parameters until the model convergence conditions are met, thereby obtaining the target remote sensing small sample change detection model.

10. A server system, characterized in that: The method comprises a server, wherein the server is configured to execute the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Small sample change detection method based on multi-scale feature extraction

    CN112668494A

  • Remote sensing image change detection method based on fine tuning CLIP

    CN119649062A

  • Receiver for communications satellite down-link reception

    US4956864A

  • Geosynchronization of an aerial image using localizing multiple features

    WO2024042508A1