Remote sensing image change detection and updating method based on multi-agent reinforcement learning

By employing multi-agent reinforcement learning technology, the problems of inaccurate feature capture and insufficient collaborative decision-making in remote sensing image change detection are solved, achieving efficient and accurate change detection and adaptive updating, thereby improving the detection accuracy and timeliness of remote sensing image data updates.

CN121746901APending Publication Date: 2026-03-27NANJING WANBO GEOGRAPHIC INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-12
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Traditional remote sensing image change detection methods suffer from several drawbacks when processing dual-temporal or multi-temporal remote sensing images. These include inaccurate capture of deep features in local areas, incomplete processing of large-size images, low efficiency, lack of efficient distributed collaborative mechanisms, insufficient communication and joint decision-making among agents, difficulty in balancing detection accuracy, data processing volume, and energy consumption, and inflexible automatic updates, all of which affect the reliability and timeliness of detection and updates.

Method used

We employ multi-agent reinforcement learning technology to perform preprocessing, local feature extraction, agent communication and decision-making on remote sensing images. We combine convolutional neural networks and the Qmix algorithm for joint decision-making, and construct a reward function through reinforcement learning training to optimize detection accuracy and energy consumption. We also set an adaptive update threshold to achieve automatic updates.

Benefits of technology

It improves the accuracy of changing area identification, optimizes detection precision and energy consumption, ensures the timeliness and consistency of remote sensing image data, verifies the reliability of detection results through multi-dimensional evaluation indicators, and adapts to the resource constraints of different orbital intelligent agents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121746901A_ABST
    Figure CN121746901A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of remote sensing images, and discloses a remote sensing image change detection and updating method based on multi-agent reinforcement learning, and the key points of the technical scheme are as follows: remote sensing image preprocessing, local feature extraction, multi-agent deployment and communication, reinforcement learning training, change detection and self-adaptive automatic updating. Through a multi-agent reinforcement learning technology, preprocessing, local feature extraction, agent communication decision and reinforcement learning training are carried out on a dual-temporal or multi-temporal remote sensing image, accurate change detection and self-adaptive automatic updating are realized, and meanwhile, the detection precision, the data processing amount and the energy consumption are balanced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of remote sensing images, more particularly, it relates to a remote sensing image change detection and updating method based on multi-agent reinforcement learning. BACKGROUND

[0002] In the technical field of remote sensing image change detection and automatic updating, the traditional method has obvious limitations when processing double-time or multi-time remote sensing images: after pre-processing, the traditional method relies on a single model to extract features, which is not accurate enough for capturing deep local features, and it is also prone to problems such as incomplete coverage and low efficiency in processing large-size images; there is a lack of efficient distributed collaborative mechanism, making it difficult for agents to effectively communicate and jointly decide, resulting in insufficient collaborative precision of change detection; at the same time, it is difficult to balance the relationship between change detection precision, data processing amount and energy consumption, the automatic updating triggering mechanism is not flexible enough, and it cannot adaptively adjust according to the difference between the detection results and historical data, ultimately affecting the reliability of change detection and the timeliness of updating.

[0003] Therefore, the present application provides a remote sensing image change detection and updating method based on multi-agent reinforcement learning, which improves the above technical problems. SUMMARY

[0004] The embodiments of the present application aim to overcome the deficiencies of the prior art, and provide a remote sensing image change detection and updating method based on multi-agent reinforcement learning. The present application uses multi-agent reinforcement learning technology to pre-process, extract local features, communicate and make decisions among agents, and train reinforcement learning for double-time or multi-time remote sensing images, achieving accurate change detection and adaptive automatic updating, while balancing detection precision, data processing amount and energy consumption.

[0005] The above technical purposes of the present application are achieved by the following technical solutions: a remote sensing image change detection and updating method based on multi-agent reinforcement learning, comprising the following steps:

[0006] S1, remote sensing image preprocessing: performing radiation correction, geometric correction and image registration on double-time or multi-time input remote sensing images to obtain standardized image data;

[0007] S2, local feature extraction: using a convolutional neural network (CNN) to extract deep local features from the standardized image data to obtain an image feature map;

[0008] S3, multi-agent deployment and communication: constructing a distributed multi-agent network, adaptively adjusting the number of agents and the observation range according to the size of the standardized image data, each agent interacting with feature information through a pre-set communication mechanism, and jointly deciding based on Q mix algorithm;

[0009] S4, reinforcement learning training: a reward function is constructed with change detection accuracy, data processing amount and energy consumption as optimization objectives, a multi-agent network is trained through gradient update, and a mutual information maximization strategy is introduced to optimize feature matching effect;

[0010] S5, change detection: based on the output of the trained multi-agent network, the change detection result is verified in multiple dimensions by using Precision, Recall, F1 score, Kappa coefficient and overall accuracy OA;

[0011] S6, adaptive automatic update: set the Kappa coefficient update threshold, when the Kappa coefficient of the verified change detection result and the historical remote sensing image data is lower than the threshold, automatically trigger the remote sensing image data update.

[0012] As a preferred technical solution of the application, the convolutional neural network in S2 is ResNet, U-Net or its improved network, which is used to capture deep semantic features and texture features at pixel level and target level in the image.

[0013] As a preferred technical solution of the application, the number of agents in S3 is positively related to the size of the image pixels, the observation range covers the local area of the image, and there is a preset overlap degree between the observation ranges of adjacent agents. The communication mechanism is a feature weighted interaction based on attention mechanism.

[0014] As a preferred technical solution of the application, the reward function in S4 is a weighted sum function, wherein the weights of change detection accuracy, data processing amount and energy consumption are dynamically adjusted according to actual application scenarios, and the sum of the weights is 1.

[0015] As a preferred technical solution of the application, the Kappa coefficient update threshold in S6 is in the range of 0.6-0.8, which can be adaptively adjusted according to the remote sensing image application scenario (such as land use monitoring, disaster emergency response).

[0016] As a preferred technical solution of the application, the input remote sensing image includes SAR remote sensing image and optical remote sensing image, and the preprocessing in S1 further includes denoising processing for SAR image and atmospheric correction processing for optical image.

[0017] As a preferred technical solution of the application, the multi-agent network in S3 is deployed on a MEO / LEO orbit agent platform, and distributed collaborative computing is used to realize parallel processing of large-size remote sensing images.

[0018] As a preferred technical solution of the application, in the reinforcement learning training process in S4, the experience replay mechanism is used to store the interaction experience of the agent, and the training stability and generalization ability are improved through random sampling.

[0019] In summary, the present application has the following beneficial effects:

[0020] Firstly, deep local features are extracted by CNN, combined with multi-agent communication and Q mix Joint decision making, combined with mutual information maximization strategy, improves the accuracy of change area recognition, F1 score and Kappa coefficient are better. At the same time, the distributed multi-agent network can adaptively adjust the number and observation range according to the image size, avoid the efficiency bottleneck of single model processing large size image, and compatible with the differentiated processing needs of SAR and optical image.

[0021] Secondly, the reward function of reinforcement learning comprehensively considers detection accuracy, data processing amount and energy consumption, optimizes network parameters through gradient update, reduces total energy consumption while ensuring performance, and adapts to resource constraints of different orbit agents (MEO / LEO). At the same time, based on Kappa coefficient, set adaptive update threshold, when the difference between new detection result and historical data reaches threshold, automatically trigger update, ensure the timeliness and consistency of remote sensing image data.

[0022] Thirdly, multi-dimensional evaluation indexes such as Precision, Recall, F1, Kappa and OA are used to verify the detection results comprehensively, reduce false positive and false negative cases, and improve the practicability and reliability of the technical scheme. BRIEF DESCRIPTION OF DRAWINGS

[0023] Figure 1 The flow chart of the remote sensing image change detection and update method based on multi-agent reinforcement learning provided by the embodiments of the present application. DETAILED DESCRIPTION

[0024] The present application will be described in detail below with specific embodiments. The following examples will help those skilled in the art to further understand the present application, but do not limit the present application in any form. It should be pointed out that, for those skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made. These all belong to the protection scope of the present application.

[0025] In order to make the purpose, technical scheme and advantages of the present application more clear and obvious, the present application will be further described in detail below combined with the drawings and examples. It should be understood that the specific embodiments described here are only used to explain the present application, and are not used to limit the present application.

[0026] Unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The term "and / or" as used in this specification includes any and all combinations of one or more of the associated listed items.

[0027] Furthermore, the technical features involved in the various embodiments of this application described below can be combined with each other as long as they do not conflict with each other.

[0028] Please refer to Figure 1 , Figure 1 A flowchart illustrating the vehicle repair shop operation evaluation method according to an embodiment of this disclosure is shown. The overall process mainly includes the following 7 steps:

[0029] Step 1: Remote Sensing Image Preprocessing

[0030] Preprocessing of dual-temporal or multi-temporal remote sensing images includes geometric registration, radiometric correction, denoising, and data standardization to ensure spatial consistency and data reliability.

[0031] Geometric registration: The RANSAC algorithm is used to correct image distortion and align the pixel coordinates of images from different time phases. The registration error satisfies Δx≤1 pixel and Δy≤1 pixel.

[0032] Radiometric correction: The digital value is converted to ground reflectance using absolute radiometric correction, as shown in the following formula:

[0033]

[0034] Where ρ is the ground reflectance, DN is the original digital value of the image, and G gain B is the gain coefficient. bias L is the bias coefficient. sun θ represents solar irradiance. sun This is the solar altitude angle.

[0035] Denoising: SAR images are despecced using a Leesigma filter, while optical images are despecced using a nonlocal mean filter to remove Gaussian noise. The filtering formula is as follows:

[0036] I denoised (x,y)=Σ (u,v)∈Ω w(x,y,u,v)×I(x+u,y+v)

[0037] Where Ω is the filter window and w(x,y,u,v) is the weight coefficient, satisfying ∑w(x,y,u,v)=1.

[0038] Step 2: Multi-agent system initialization

[0039] A distributed multi-agent network is constructed, where each agent is responsible for local area observation and decision-making in remote sensing imagery. Specific initialization parameters include:

[0040] Number of agents N: Adaptively set according to image size to meet the following requirements. (W and H are the image width and height, and ω is the size of the agent's observation window);

[0041] Observation window: The default setting is ω×ω=20×20 pixels, and the agent's initial position is p. i (0) Generated using a uniform random distribution;

[0042] Action space: The action set A for each agent. i = {up, down, left, right}, step size equals the observation window size, action constraint is: if the agent reaches the image boundary, the action is invalid and the position remains unchanged;

[0043] State space: Agent states Where b i (t) represents a local feature. The average communication message λ for other agents i (t) is the position code, E i (t) represents the current energy consumption.

[0044] Step 3: Local Feature Extraction

[0045] Each agent extracts deep features of the local observation region through a convolutional neural network (CNN). The CNN structure includes three convolutional layers and two max-pooling layers. The feature extraction process is as follows:

[0046] Local observation generation: The local observation of agent i at time t is as follows:

[0047] o i (t)=O(I,p i (t),w)

[0048] Where I represents the preprocessed remote sensing image, and p i (t)=(x i (t), y i (t) represents the position coordinates of the agent.

[0049] Convolutional feature extraction: Multi-scale features are extracted through convolution operations. The output of the convolution at the 1st layer is:

[0050]

[0051] in, For convolution kernel weights, For bias terms, is the convolution operator, and ReLU is the activation function.

[0052] Position encoding: The agent's position is encoded through a fully connected layer, and the CeLU activation function is used to avoid gradient vanishing.

[0053] λ i (t)=CeLU(W p ·p i (t)+b p )

[0054] CeLU(x)=max(0,x)+min(0,exp(x)-1)

[0055] Among them, W p b p These are the parameters for the location coding layer.

[0056] Step 4: Multi-agent communication and decision-making

[0057] Based on gated cyclic unit (GRU) and Q mix Networks enable communication and joint decision-making among intelligent agents, specifically including:

[0058] Communication message generation: The communication messages of each agent are generated from the hidden states of the predictive GRU and the decision GRU.

[0059]

[0060] Among them, h i (t) represents the predicted GRU hidden state. The decision-making GRU hidden state is represented by θ7, which is a parameter of the communication module.

[0061] Average message passing: Agent i receives the average number of messages from other agents:

[0062]

[0063] GRU State Update: The state update formulas for Predictive GRU and Decision GRU are as follows:

[0064] r(t)=σ(W r ·[h(t),x(t)])

[0065] z(t)=σ(W z ·[h(t), x(t)])

[0066]

[0067] Where r(t) is the reset gate, z(t) is the update gate, σ is the Sigmoid function, and ⊙ is the element-wise product;

[0068] Q mix Joint decision-making: combining the local Q-values ​​of each agent. i (s i (t), a i (t) is fused into a global Q-value through a hybrid network:

[0069] Q tot (s(t), a(t))=mix(Q1,Q2,...,Q N ;θ mix )

[0070] Where s(t)={s1(t),s2(t),...,s N Let a(t) be the global state, and a(t) = {a1(t), a2(t), ..., a2(t)}. N (t)} represents a joint action, θ mix These are mixed network parameters.

[0071] Step 5: Reinforcement Learning Training

[0072] The training process is modeled based on a partially observable Markov decision process (POMDP), with the goal of maximizing the cumulative reward while satisfying energy constraints.

[0073] Maximizing mutual information: Improve coordination accuracy by maximizing the mutual information between the agent's trajectory / action perception and the actual trajectory / action.

[0074]

[0075]

[0076] in, Let i be the trajectory perception of agent j. For motion perception, C j For agent identification, the lower bound of mutual information is achieved by minimizing the KL divergence:

[0077]

[0078] Reward function design: comprehensively considering change detection accuracy, data processing volume, and energy consumption.

[0079] r(t) = α·CD acc (t)+β·D proc (t)-γ·E total (t)

[0080] Among them: CD acc(t) represents the change detection accuracy, calculated using the F1 score:

[0081]

[0082] Wherein: TP indicates a true positive, FP indicates a false positive, and FN indicates a false negative.

[0083] D proc (t) represents the data processing volume. For SAR images, the data is calculated using microwave signal reception volume, and for optical images, it is calculated using pixel information entropy.

[0084] Total energy consumption; MEO agent energy consumption:

[0085]

[0086] LEO agent energy consumption:

[0087]

[0088] Among them: A i,j (t) represents the amount of LEO data processed by MEO, C M For MEO computation rate, P M P L P T These represent MEO calculated power, LEO calculated power, and transmission power, respectively, H. i,j (t) is an access control variable;

[0089] α, β, and γ are weighting coefficients that satisfy α + β + γ = 1.

[0090] Gradient Update: The network parameters Θ = [θ1, θ2, ..., θ7] are updated using stochastic gradient descent. The gradient of the objective function is:

[0091]

[0092] Where C is the number of trajectory samples, p k Let r be the probability of the k-th trajectory occurring. k For trajectory rewards.

[0093] Step 6: Change Detection and Automatic Updates

[0094] Change Map Generation: Based on the joint decision-making results of multiple agents, a binary change map (CM) is generated using the Softmax function.

[0095]

[0096] Where, q i =f4(h i(T); θ4) represents the local prediction result of agent i, and T is the training time step;

[0097] Automatic update mechanism: Set an update threshold δ. When the Kappa coefficient of the new detection result and the historical result satisfies κ < δ, automatic update is triggered.

[0098]

[0099] TN stands for true negative, PCC for correct classification percentage, and PRE for expected consistency rate.

[0100] Step 7: Use Precision, Recall, F1, Kappa, and Overall Accuracy (OA) as evaluation metrics to comprehensively verify the reliability of the change detection results:

[0101]

[0102] Example:

[0103] Number of agents N = 10 (can be adjusted according to image size); observation window size w = 20 pixels; training time step T = 5; trajectory sampling times C = 3; learning rate η = 0.002; weight coefficients α = 0.5, β = 0.3, γ = 0.2; update threshold δ = 0.85; Qmix network hidden layer dimension: 256; GRU hidden layer dimension: 128.

[0104] Data input: Input dual-temporal SAR image (15m resolution) and multispectral image (10m resolution), with an image size of 1024×1024 pixels;

[0105] Preprocessing: Geometric registration was performed using the RANSAC algorithm, radiometric correction was converted to ground reflectance, SAR images were denoised using Leesigma filtering, and multispectral images were denoised using nonlocal mean filtering.

[0106] Agent initialization: Generate 10 agents, with their initial positions evenly distributed and their action space set to move up, down, left, and right;

[0107] Feature extraction: Each agent extracts 20×20 local region features through CNN, and the position encoding uses the CeLU activation function;

[0108] Communication and Decision-Making: Communication messages are generated based on GRU, and local Q-values ​​are fused through a Qmix network to obtain joint decision results;

[0109] Reinforcement learning training: Iterative training for 200 epochs, sampling 3 trajectories per epoch, and updating parameters through stochastic gradient descent;

[0110] Change detection: Generate binary change maps, detect building construction changes using SAR images, and detect vegetation cover changes using multispectral images;

[0111] Automatic update: Calculate the Kappa coefficient between the new result and the historical results. If κ = 0.82 < 0.85, trigger automatic update.

[0112] Results output: Change detection map and evaluation indicators are output. The F1 score is 0.987, the OA score is 0.991, and the Kappa score is 0.978.

[0113] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A method for remote sensing image change detection and updating based on multi-agent reinforcement learning, characterized in that, The method includes the following steps: S1. Remote sensing image preprocessing: Perform radiometric correction, geometric correction and image registration on dual-temporal or multi-temporal input remote sensing images to obtain standardized image data. S2. Local Feature Extraction: A convolutional neural network (CNN) is used to perform deep local feature extraction on standardized image data to obtain image feature maps. S3. Multi-agent Deployment and Communication: Construct a distributed multi-agent network, adaptively adjusting the number of agents and observation range based on the size of standardized image data. Each agent interacts with feature information through a pre-defined communication mechanism, and based on... The algorithm performs joint decision-making; S4. Reinforcement learning training: Construct a reward function with change detection accuracy, data processing volume and energy consumption as optimization objectives, and perform reinforcement learning training on the multi-agent network through gradient update. At the same time, introduce a mutual information maximization strategy to optimize the feature matching effect. S5. Change Detection: Based on the change detection results output by the trained multi-agent network, multi-dimensional verification is performed using Precision, Recall, F1 score, Kappa coefficient, and overall accuracy OA. S6. Adaptive Automatic Update: Set a Kappa coefficient update threshold. When the Kappa coefficient of the verified change detection result and the historical remote sensing image data are lower than the threshold, the remote sensing image data update is automatically triggered.

2. The remote sensing image change detection and update method based on multi-agent reinforcement learning according to claim 1, characterized in that, The convolutional neural network described in S2 is ResNet, U-Net, or an improved version thereof, used to capture pixel-level and target-level deep semantic and texture features in images.

3. The remote sensing image change detection and update method based on multi-agent reinforcement learning according to claim 1, characterized in that, The number of agents in S3 is proportionally distributed in a positive correlation with the image pixel size. The observation range covers a local area of ​​the image and there is a preset overlap between the observation ranges of adjacent agents. The communication mechanism is a feature-weighted interaction based on an attention mechanism.

4. The remote sensing image change detection and update method based on multi-agent reinforcement learning according to claim 1, characterized in that, The reward function described in S4 is a weighted summation function, in which the change detection accuracy weight, data processing volume weight, and energy consumption weight are dynamically adjusted according to the actual application scenario, and the sum of the weights is 1.

5. The remote sensing image change detection and update method based on multi-agent reinforcement learning according to claim 1, characterized in that, The Kappa coefficient update threshold mentioned in S6 has a value range of 0.6-0.8, which can be adaptively adjusted according to the application scenario of remote sensing imagery (such as land use monitoring and disaster emergency response).

6. The remote sensing image change detection and update method based on multi-agent reinforcement learning according to claim 1, characterized in that, The input remote sensing images include SAR remote sensing images and optical remote sensing images. The preprocessing in S1 also includes denoising processing for SAR images and atmospheric correction processing for optical images.

7. The remote sensing image change detection and update method based on multi-agent reinforcement learning according to claim 1, characterized in that, In S3, the multi-agent network is deployed on the MEO / LEO orbital agent platform, enabling parallel processing of large-size remote sensing images through distributed collaborative computing.

8. The remote sensing image change detection and update method based on multi-agent reinforcement learning according to claim 1, characterized in that, In the reinforcement learning training process in S4, an experience replay mechanism is used to store the agent's interaction experience, and random sampling is used to improve training stability and generalization ability.