Anti-Hallucination Multimodal Large Model Change Detection Method for Autonomous Land Management Agent
Through the anti-illusion multimodal large-modal model change detection method of autonomous land management agents, the complex or dynamic changing environment is solved, and the problem of insufficient robustness and generalization in the prior art is achieved, and the accuracy and reliability of change detection are achieved.
Patent Information
- Application Number
- CN202510280552.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-03-11
AI Technical Summary
The prior art faces the problem of insufficient robustness and generalization when dealing with complex or dynamically changing environments, especially in scenarios that deviate from the distribution of training data or highly variable environmental conditions.
The anti-illusion multimodal large-modal model change detection method of autonomous land management agent is adopted. By receiving remote sensing images at two different time points, a normalized time difference map is generated using the SigLip image encoder, and a large language model is imported into the process with text markers, iterative updates and convergence processing is performed, and the change detection map is finally generated and predicted aggregation is performed.
It significantly improves the generalization ability of the model in complex and dynamic environments, reduces hallucinations, improves the accuracy and reliability of change detection, and ensures the robustness and consistency of detection results.
Smart Images

Figure CN119810673B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of change detection, and particularly relates to a change detection method for an anti-hallucination multimodal large model of an autonomous land management agent. Background Art
[0002] Multimodal data processing refers to the fusion and analysis of information from different modalities (such as images and texts) to achieve precise detection of target changes. Change Detection (CD) is one of the core tasks of remote sensing technology, aiming to identify changes in the earth's surface at different time periods by analyzing multi-temporal remote sensing images.
[0003] Currently, change detection (CD) methods based on deep learning (DL) are mainly divided into two categories: methods based on convolutional neural networks (CNNs) and methods based on the Transformer architecture.
[0004] CNN-based CD methods perform well in extracting local spatial features and are suitable for learning complex features from high-resolution satellite and aerial images. These methods usually combine preprocessing techniques such as image alignment and segmentation to enhance the model's change detection ability. However, due to the inherent local receptive field limitation of CNNs, they perform poorly in capturing global context and long-frequency information. In addition, when the data volume is limited, CNN methods are prone to overfitting problems and are difficult to effectively adapt to dynamic or subtle changes, especially in large-scale dataset or multi-temporal image processing scenarios.
[0005] Methods based on the Transformer architecture were originally designed for natural language processing (NLP) tasks and have gradually attracted attention in the CD field in recent years because their attention mechanism can model long-range dependencies and global context information. Transformer-based CD methods improve the accuracy of change detection by capturing spatio-temporal relationships and are particularly suitable for complex scenarios and dynamic environments. However, these methods have limitations in extracting local features, which is crucial for detecting fine-grained changes in high-resolution images. In addition, the computational complexity of Transformer is relatively high, requiring a large amount of hardware resources.
[0006] Therefore, although the existing technologies have made significant improvements in performance, they still face challenges when dealing with complex or dynamic change environments. For example, when encountering scenarios that deviate from the training data distribution, or when the detection task involves highly variable environmental conditions, these methods often show poor robustness and generalization. Summary of the Invention
[0007] The purpose of the embodiments of the present invention is to provide a change detection method for an anti-hallucination multimodal large model of an autonomous land management agent, aiming to solve the problems raised in the background art.
[0008] To achieve the above object, the embodiments of the present invention provide the following technical solutions:
[0009] A method for change detection of an anti-hallucination multimodal large model of an autonomous land management agent, the method specifically comprising the following steps:
[0010] Receive remote sensing images at two different time points, and capture the differences between the two remote sensing images through a SigLip image encoder to generate a normalized time difference map;
[0011] Obtain text tokens, combine the normalized time difference map with the text tokens, and import them into a preset large language model for processing to generate a rough change map;
[0012] Based on the normalized time difference map, the text tokens, and the previous calibration mask, construct an iterative input, perform iterative update and convergence processing, and obtain a change detection map;
[0013] Receive multiple change detection tasks, generate multiple inference paths, and perform prediction aggregation to generate a final prediction.
[0014] As a further limitation of the technical solution of the embodiment of the present invention, the step of receiving remote sensing images at two different time points, and capturing the differences between the two remote sensing images through a SigLip image encoder to generate a normalized time difference map specifically comprises the following steps:
[0015] Receive remote sensing images at two different time points;
[0016] Through a SigLip image encoder, input the two remote sensing images into a shared encoder network for feature representation to obtain two feature maps;
[0017] Through a projection head, project the two feature maps into a shared latent space to obtain two latent space dimensions;
[0018] Calculate a normalized time difference map according to the two latent space dimensions.
[0019] As a further limitation of the technical solution of the embodiment of the present invention, the expressions of the two feature maps are:
[0020] ;
[0021] ;
[0022] Wherein, and are the feature maps corresponding to the two remote sensing images respectively, and represent two remote sensing images, denote the encoder network, is the set of parameters of the encoder network;
[0023] The expressions for the two latent space dimensions are:
[0024] ;
[0025] ;
[0026] where, and are the latent space dimensions corresponding to the two remote sensing images respectively, denote the projection head, is the set of parameters of the projection head;
[0027] The calculation formula of the normalized time difference map is:
[0028] ;
[0029] where, is a small constant introduced for numerical stability.
[0030] As a further limitation of the technical solution of the embodiment of the present invention, the obtaining of the text marker, combining the normalized time difference map with the text marker, and importing it into a preset large language model for processing to generate a rough change map specifically includes the following steps:
[0031] Obtain the text marker;
[0032] Combine the normalized time difference map with the text marker to generate a multimodal input;
[0033] Import the multimodal input into a preset large language model for processing to generate a rough change map.
[0034] As a further limitation of the technical solution of the embodiment of the present invention, the expression of the multimodal input is:
[0035] ;
[0036] where, is the multimodal input, is the text marker;
[0037] The expression of the rough change map is:
[0038] ;
[0039] where, represents the rough change map.
[0040] As a further limitation of the technical solution of the embodiment of the present invention, constructing an iterative input based on the normalized time difference map, the text markers, and the previous calibration mask, and performing iterative update and convergence processing to obtain a change detection map specifically includes the following steps:
[0041] Construct an iterative input based on the normalized time difference map, the text markers, and the previous calibration mask;
[0042] Perform iterative update of the change map according to the iterative input, and calculate the current calibration mask;
[0043] The iterative process continues until the change map converges, and finally outputs the change detection map.
[0044] As a further limitation of the technical solution of the embodiment of the present invention, the expression for iterative update of the change map is:
[0045] ;
[0046] where is the change map for iterative update, is the iterative input;
[0047] The calculation formula for the current calibration mask is:
[0048] ;
[0049] The expression for the iterative process to continue until the change map converges is:
[0050] ;
[0051] where is a preset convergence threshold.
[0052] As a further limitation of the technical solution of the embodiment of the present invention, receiving multiple change detection tasks, generating multiple inference paths, and performing prediction aggregation to generate a final prediction specifically includes the following steps:
[0053] Receive multiple change detection tasks;
[0054] Generate multiple inference paths according to the multiple change detection tasks, and associate multiple corresponding logical assumptions or conditions;
[0055] Generate intermediate CD maps corresponding to the multiple inference paths;
[0056] Aggregate and calculate the final CD map based on the multiple intermediate CD maps to generate a final prediction.
[0057] As a further limitation of the technical solution of the embodiment of the present invention, the expression for generating multiple inference paths according to the multiple change detection tasks and associating multiple corresponding logical hypotheses or conditions is:
[0058] ;
[0059] wherein, is the th inference path, is the logical hypothesis or condition corresponding to the th inference path;
[0060] The expressions for the multiple intermediate CD diagrams are:
[0061] ;
[0062] wherein, is the intermediate CD diagram corresponding to the th inference path;
[0063] The expression for the final CD diagram is:
[0064] ;
[0065] wherein, is the final CD diagram, represents an aggregation function.
[0066] As a further limitation of the technical solution of the embodiment of the present invention, another expression for the final CD diagram is:
[0067] ;
[0068] wherein, is the path confidence weight corresponding to the th inference path.
[0069] Compared with the prior art, the beneficial effects of the present invention are:
[0070] In the embodiments of the present invention, by receiving remote sensing images at two different time points, capturing the differences between the two remote sensing images to generate a Normalized Time Difference Map; importing it into a preset large language model for processing to generate a rough change map; constructing an iterative input based on the Normalized Time Difference Map, text markers, and the previous calibration mask, performing iterative update and convergence processing to obtain a change detection map; receiving multiple change detection tasks, generating multiple inference paths, and performing prediction aggregation to generate a final prediction. It can improve the generalization ability of the model in complex and dynamic environments, gradually optimize the detection results, significantly reduce the hallucination problem, and improve the accuracy and reliability of change detection, ensuring the robustness and consistency of the detection results, and making it applicable to complex and changing remote sensing scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention.
[0072] Figure 1 The flowchart of the anti-hallucination multimodal large model change detection method for the autonomous land management agent provided by the embodiments of the present invention is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0073] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the following further details the present invention with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0074] It can be understood that in the prior art, for the CNN-based CD method, due to the inherent local receptive field limitation of CNN, it performs poorly in capturing global context and long-frequency information. In addition, when the data volume is limited, the CNN method is prone to overfitting problems and is difficult to effectively adapt to dynamic or subtle changes, especially in large-scale dataset or multi-temporal image processing scenarios; for the method based on the Transformer architecture, there are limitations in extracting local features, which is crucial for detecting fine-grained changes in high-resolution images. In addition, the computational complexity of the Transformer is relatively high, requiring a large amount of hardware resources. Therefore, although the prior art has made significant improvements in performance, it still faces challenges when dealing with complex or dynamic change environments. For example, when encountering scenarios that deviate from the training data distribution, or when the detection task involves highly variable environmental conditions, these methods often exhibit poor robustness and generalization.
[0075] To solve the above problems, in the embodiments of the present invention, by receiving remote sensing images at two different time points, and through the SigLip image encoder, capturing the differences between the two remote sensing images to generate a normalized time difference map; obtaining text tokens, combining the normalized time difference map with the text tokens, and importing them into a preset large language model for processing to generate a rough change map; based on the normalized time difference map, text tokens, and the previous calibration mask, constructing an iterative input, performing iterative update and convergence processing to obtain a change detection map; receiving multiple change detection tasks, generating multiple inference paths, and performing prediction aggregation to generate a final prediction. It can improve the generalization ability of the model in complex and dynamic environments, gradually optimize the detection results, significantly reduce the hallucination problem, and improve the accuracy and reliability of change detection, ensuring the robustness and consistency of the detection results, and making it applicable to complex and changeable remote sensing scenarios.
[0076] Figure 1 Fig. shows the flowchart of the anti-hallucination multimodal large model change detection method for the autonomous land management agent provided by the embodiments of the present invention.
[0077] Specifically, for the anti-hallucination multimodal large model change detection method of the autonomous land management agent, the method specifically includes the following steps:
[0078] Step S101, receive remote sensing images at two different time points, and through the SigLip image encoder, capture the differences between the two remote sensing images to generate a normalized time difference map.
[0079] In the embodiments of the present invention, by receiving remote sensing images at two different time points of the autonomous land management agent, using the SigLip image encoder, inputting the two remote sensing images into a shared encoder network for feature representation to obtain two feature maps, and then through a projection head, projecting the two feature maps into a shared latent space to obtain two latent space dimensions, and further calculating the normalized time difference map according to the two latent space dimensions. Specifically, the expressions of the two feature maps are:
[0080] ;
[0081] ;
[0082] Among them, and are respectively the feature maps corresponding to the two remote sensing images, and represent the two remote sensing images, represents the encoder network, is the parameter set of the encoder network;
[0083] The expression of the two latent space dimensions is:
[0084] ;
[0085] ;
[0086] Among them, and are the potential space dimensions corresponding to two remote sensing images respectively, represents the projection head, is the parameter set of the projection head;
[0087] The calculation formula of the normalized time difference map is:
[0088] ;
[0089] Among them, is a small constant introduced for numerical stability.
[0090] It can be understood that the SigLip encoder processes images from two different time points to generate feature embeddings representing the time difference of the images. These feature representations are input into the LLM together with the corresponding text data, and the LLM generates preliminary change predictions by capturing patterns in the visual and context modalities. As a multimodal model, the LLM utilizes the complementarity of image and text features to enhance the understanding of changes in the environment, thereby being able to capture complex relationships that are difficult to handle by traditional CNN or Transformer models, especially in novel or unexpected change scenarios.
[0091] It can be understood that the normalized time difference map captures the magnitude and direction of time changes in the latent feature space, capable of highlighting key changes between images, such as urbanization or deforestation, while suppressing background noise.
[0092] Step S102, obtain text tokens, combine the normalized time difference map with the text tokens, and import them into a preset large language model for processing to generate a rough change map.
[0093] In the embodiments of the present invention, by obtaining text tokens, combining the normalized time difference map with the text tokens to generate multimodal input, and then importing the multimodal input into a preset large language model for processing to generate a rough change map. Specifically, the expression of the multimodal input is:
[0094] ;
[0095] Among them, is the multimodal input, is the text token;
[0096] The expression of the rough change map is:
[0097] ;
[0098] Among them, represents a rough change map.
[0099] It can be understood that in the multi-modal framework, the SigLip encoder processes each image separately to generate a feature representation, emphasizing temporal differences rather than static attributes. The generated normalized temporal difference map is combined with text tokens and used as the main input for the subsequent stages of the model (including the CVA calibration loop), thereby guiding the model to focus on detecting meaningful changes rather than irrelevant background features. The modular and robust difference extraction enables the system to maintain accuracy and adaptability in various environments.
[0100] It can be understood that the rough change map provides a preliminary estimate of the change probability for each pixel. However, this rough prediction usually contains spatial inaccuracies and needs to be further corrected through an iterative optimization process.
[0101] Step S103: Based on the normalized temporal difference map, the text tokens, and the previous calibration mask, construct an iterative input, perform iterative update and convergence processing, and obtain a change detection map.
[0102] In the embodiment of the present invention, based on the normalized temporal difference map, text tokens, and the previous calibration mask, an iterative input is constructed, and then, based on the iterative input, the change map is iteratively updated, and the current calibration mask is calculated. The iterative process continues until the change map converges, and finally, a change detection map is output. Specifically, the expression for the iterative update of the change map is:
[0103] ;
[0104] Among them, is the change map for iterative update, is the iterative input;
[0105] The calculation formula for the current calibration mask is:
[0106] ;
[0107] The expression for the iterative process to continue until the change map converges is:
[0108] ;
[0109] Among them, is a preset convergence threshold.
[0110] It is understandable that the calibration mask is used to correct the predicted change map, ensuring that the optimization process focuses on meaningful temporal changes while reducing the impact of noise; the change detection map combines the semantic understanding of the LLM with the spatial accuracy of the CVA, reducing hallucination phenomena and improving detection accuracy.
[0111] Step S104: Receive multiple change detection tasks, generate multiple inference paths, and perform prediction aggregation to generate a final prediction.
[0112] In the embodiment of the present invention, by receiving multiple change detection tasks, multiple inference paths are generated according to the multiple change detection tasks, and multiple corresponding logical assumptions or conditions are associated. Then, intermediate CD maps corresponding to the multiple inference paths are generated. Furthermore, based on the multiple intermediate CD maps, the final CD map is calculated by aggregation to generate a final prediction. Specifically, the expression for generating multiple inference paths and associating multiple corresponding logical assumptions or conditions according to the multiple change detection tasks is:
[0113] ;
[0114] where is the th inference path, is the logical assumption or condition corresponding to the th inference path;
[0115] The expression for the multiple intermediate CD maps is:
[0116] ;
[0117] where is the intermediate CD map corresponding to the th inference path;
[0118] The expression for the final CD map is:
[0119] ;
[0120] where is the final CD map, represents an aggregation function;
[0121] Furthermore, another expression for the final CD map is:
[0122] ;
[0123] where is the path confidence weight corresponding to the th inference path.
[0124] It can be understood that step S104 is a CoT reasoning process, which effectively addresses the uncertainties in remote sensing data, such as noise or ambiguity in observational changes. By generating multiple hypotheses and evaluating these hypotheses through aggregation, the model ensures that the most consistent and credible explanations are preferentially selected. For example, if one path attributes the detected change to urbanization, while another path suggests it is a seasonal vegetation change, the aggregation step can synthesize these hypotheses, thereby reducing the risk of overfitting to irrelevant features. This ability enables the model to consider the complex environmental factors affecting the observed changes.
[0125] It can be understood that in addition to handling uncertainties, CoT reasoning significantly enhances the robustness of the model. By simultaneously exploring multiple hypotheses, the model reduces the likelihood of misinterpreting background noise or pseudo-patterns as real changes. Each reasoning path aims to explore different hypotheses or explanations, refining the predictions in intermediate steps, enabling the model to better generalize to new or unseen environments. This is particularly important in remote sensing as data quality can vary due to atmospheric interference, sensor limitations, or uneven lighting conditions.
[0126] It can be understood that another key advantage of CoT reasoning is its ability to reduce hallucination phenomena, that is, false alarms or irrelevant changes caused by overgeneralization or noise. By generating multiple reasoning paths, CoT reasoning reduces the impact of hallucinations that any single path might trigger. The aggregation process ensures that such false predictions are overridden by the more consistent and credible path results, thereby ultimately improving the overall accuracy of the model. This approach is particularly valuable in complex tasks such as urbanization detection, where the model must distinguish real changes (such as new buildings) from irrelevant phenomena (such as seasonal vegetation changes).
[0127] It can be understood that through iterative hypothesis generation and optimization, CoT reasoning provides a structured framework for change detection. Each reasoning path is refined through intermediate steps, incorporating additional features, semantic context, and prior knowledge to improve the predictions. The final CD output is obtained by aggregating the results of these paths, giving more weight to the most consistent and credible hypotheses. This iterative and structured reasoning makes CoT reasoning a powerful tool for achieving reliable and accurate change detection in dynamic and noisy remote sensing environments.
[0128] It should be understood that although the steps in the flowcharts of the embodiments of the present invention are shown in sequence according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in each embodiment may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0129] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0130] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0131] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent for the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several variations and improvements can still be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the patent for the present invention shall be subject to the appended claims.
[0132] The foregoing is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. Hallucination-resistant multimodal large-scale change detection method for autonomous land management agents, characterized by: The method specifically comprises the following steps: Receiving remote sensing images at two different time points, performing difference capture on the two remote sensing images through a SigLip image encoder, and generating a normalized time difference map; Acquire text tags, combine the normalized time difference graph with the text tags, import them into a preset large language model for processing, and generate a rough change graph; Based on the normalized time difference map, the text mark and the previous calibration mask, construct an iterative input, perform iterative update and convergence processing, and obtain a change detection map; Receive multiple change detection tasks, generate multiple reasoning paths, and perform prediction aggregation to generate the final prediction; The current calibration mask is calculated by the CVA algorithm, and its calculation formula is: ; and Represents two remote sensing images, is the pixel coordinate.
2. The anti-hallucination multimodal large model change detection method for autonomous land management agents according to claim 1, characterized in that The receiving of two remote sensing images at different time points, performing difference capture on the two remote sensing images through a SigLip image encoder, and generating a normalized time difference map specifically comprises the following steps: Receive remote sensing images at two different time points; The two remote sensing images are input into a shared encoder network through a SigLip image encoder for feature representation, thereby obtaining two feature maps; Projecting the two feature maps into a shared latent space through a projection head to obtain two latent space dimensions; A normalized temporal difference map is computed based on two of the latent space dimensions.
3. The anti-hallucination multimodal large model change detection method for autonomous land management agents according to claim 2, characterized in that: The expressions of the two feature maps are: ; ; in, and are the feature maps corresponding to the two remote sensing images, and Represents two remote sensing images, represents the encoder network, is the parameter set of the encoder network; The expressions for the two latent space dimensions are: ; ; in, and are the latent space dimensions corresponding to the two remote sensing images, Represents the projection head, is the parameter set of the projection head; is the pixel coordinate, For The features extracted from for The features extracted from The calculation formula of the normalized time difference map is: ; in, Small constant introduced for numerical stability.
4. The hallucination-resistant multimodal large model change detection method for autonomous land management agents according to claim 3, characterized in that: The obtaining of the text mark, combining the normalized time difference map with the text mark, and importing the normalized time difference map into a preset large language model for processing, and generating a rough change map specifically comprises the following steps: Get the text tag; Combining the normalized time difference map with the text markup to generate a multimodal input; The multimodal input is imported into a preset large language model for processing to generate a rough change map.
5. The anti-hallucination multimodal large model change detection method for autonomous land management agents according to claim 4, characterized in that: The expression of the multimodal input is: ; in, For multimodal input, Mark for text; The expression of the rough change map is: ; in, Represents a rough change diagram.
6. The hallucination-resistant multimodal large model change detection method for autonomous land management agents according to claim 5, characterized in that: The step of constructing an iterative input based on the normalized time difference map, the text mark and the previous calibration mask, performing iterative updating and convergence processing, and obtaining a change detection map specifically includes the following steps: constructing an iterative input based on the normalized time difference map, the textual label and a previous calibration mask; Iteratively updating the change map according to the iterative input, and calculating a current calibration mask; The iterative process continues until the change graph converges, and finally the change detection graph is output.
7. The hallucination-resistant multimodal large model change detection method for autonomous land management agents according to claim 6, characterized in that: The expression for iterative update of the change graph is: ; in, is the change graph for iterative updates, Input for iteration; The iteration process continues until the expression of the change graph converges is: ; in, is the preset convergence threshold.
8. The hallucination-resistant multimodal large model change detection method for autonomous land management agents according to claim 5, characterized in that: The receiving of multiple change detection tasks, generating multiple reasoning paths, and performing prediction aggregation to generate a final prediction specifically includes the following steps: Receive multiple change detection tasks; According to the plurality of change detection tasks, a plurality of reasoning paths are generated, and a plurality of corresponding logical assumptions or conditions are associated; Generate intermediate CD graphs corresponding to multiple reasoning paths; Based on the multiple intermediate CD graphs, a final CD graph is aggregated and calculated to generate a final prediction.
9. The hallucination-resistant multimodal large model change detection method for autonomous land management agents according to claim 8, characterized in that: The expression of generating multiple reasoning paths according to the multiple change detection tasks and associating multiple corresponding logical assumptions or conditions is: ; in, For the The reasoning path, For the The logical assumptions or conditions corresponding to the reasoning paths; The expressions of the plurality of intermediate CD graphs are: ; in, For the The intermediate CD graph corresponding to the reasoning path; The expression of the final CD graph is: ; in, For the final CD map, Represents an aggregate function.
10. The hallucination-resistant multimodal large model change detection method for autonomous land management agents according to claim 9, characterized in that: Another expression of the final CD graph is: ; in, For the The path confidence weight corresponding to the inference path.
Citation Information
Patent Citations
Remote sensing image comparative analysis method based on multi-view fusion
CN115830448A
Systems and methods for resolving salient changes in earth observations across time
US20240362898A1