A natural resource monitoring false alarm suppression method and device based on a multi-modal large model

CN122551183APending Publication Date: 2026-08-11SURVEYING & MAPPING INST LANDS & RESOURCE DEPT OF GUANGDONG PROVINCE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-03
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0005]本发明的目的在于提供一种基于多模态大模型的自然资源监测虚警抑制方法及装置,用于解决现有技术缺乏业务语义理解导致虚警抑制不彻底,且模型判别过程缺乏可解释性的问题

Benefits of technology

[0017] Compared with existing technologies, the advantages of this invention are as follows: This invention obtains operational attribute data by performing spatial overlay analysis on the vectors of changed patches, and combines this with the cropping results of preceding and following temporal remote sensing images, using both as input data for a multimodal large model. This allows the model inference process to simultaneously incorporate image visual features and semantic information from natural resource monitoring operations. Unlike existing technologies that rely solely on image spectral and texture features for false alarm suppression, this invention can distinguish between ground feature change scenarios with similar visual features but different operational attributes, reducing misjudgments of false change patches and improving the completeness of false alarm suppression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122551183A_ABST
    Figure CN122551183A_ABST
Patent Text Reader

Abstract

This invention discloses a method and apparatus for suppressing false alarms in natural resource monitoring based on a multimodal large model, belonging to the fields of natural resource monitoring and artificial intelligence technology. This method involves performing spatial overlay analysis on historical change patch vectors and attaching business attributes; adaptively cropping remote sensing images from different time phases to generate image slices with visual guidance markers; constructing system prompts corresponding to business rules; generating initial samples based on a pre-trained teacher model and manually correcting them to obtain training samples; using efficient parameter fine-tuning to achieve domain adaptation of the multimodal large model; deploying the model in a master-slave architecture and performing multimodal fusion inference through a distributed cluster; outputting verifiable patch authenticity judgment results with logical chains; and eliminating false change patches based on the judgment results to achieve false alarm suppression. This method can be applied to scenarios of routine dynamic monitoring of natural resources and intelligent interpretation of remote sensing images, improving the semantic discrimination capability of change patches and reducing the number of false alarms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of interdisciplinary technology of remote sensing monitoring of natural resources and artificial intelligence, specifically to the field of pseudo-change suppression technology for remote sensing image change detection. Background Technology

[0002] With the development of routine natural resource monitoring, remote sensing image change detection technology has become an important technical means for identifying land cover changes and verifying land use changes. This technology extracts changed geographic units, or change patches, from remote sensing images of different time periods, providing data support for natural resource supervision. However, in practical applications, factors such as the surface environment, image imaging conditions, and the similarity of spectral characteristics of land features can lead to a large number of false change patches during the extraction process. These false alarm patches require significant manual verification and removal, becoming a major factor hindering the automation and large-scale implementation of natural resource monitoring.

[0003] Currently, mainstream false alarm suppression technologies can be broadly categorized into three types. The first type is post-processing suppression based on spectral and texture features. This type of scheme removes visual noise-like pseudo-changes in images through threshold filtering, morphological filtering, and other operations. However, it can only handle simple noise interference and has limited ability to identify complex pseudo-change scenarios in business applications. The second type is implicit suppression based on deep learning. This type of scheme uses convolutional neural networks or visual transformer change detection models to suppress pseudo-changes at the feature level through feature enhancement and contrastive learning. This type of scheme relies on statistical induction from a large number of labeled samples. The models are mostly end-to-end black-box structures, only outputting binary discrimination results or probability values, and cannot provide traceable interpretation criteria and reasoning processes, making it difficult to meet the compliance requirements of business audits. The third type is explicit filtering based on rules and multi-source data. This type of scheme removes pseudo-changes by overlaying vector rules such as land use status maps and control red lines through spatial overlay operations. For non-image text or vector rules, the application of such solutions is mostly limited to simple spatial overlay in the later stage. They cannot achieve deep integration of feature level and inference level, and it is difficult to use prior business knowledge for deep logical verification. The accuracy of the judgment results is greatly limited.

[0004] The aforementioned false alarm suppression schemes essentially still rely on statistical induction of image spectral information and texture features, lacking a deep understanding of the semantics of natural resource monitoring business. They cannot effectively distinguish between scenarios with similar visual features but significantly different business attributes, resulting in incomplete false alarm suppression and difficulty in meeting the engineering implementation needs of complex business scenarios. Summary of the Invention

[0005] The purpose of this invention is to provide a method and apparatus for suppressing false alarms in natural resource monitoring based on a multimodal large model, which solves the problems of incomplete false alarm suppression due to the lack of business semantic understanding in the existing technology, and the lack of interpretability in the model discrimination process.

[0006] To achieve the above objectives, the present invention provides the following technical solution: In a first aspect, the present invention provides a method for suppressing false alarms in natural resource monitoring based on a multimodal large model, comprising the following steps: Acquire vectors of changes in natural resource monitoring patches and their corresponding preceding and following temporal remote sensing images; Perform spatial overlay analysis on the changed patch vectors to obtain target patch data for the attached service attributes; Spatial matching and cropping processing is performed on the target patch data and the preceding and following temporal remote sensing images to obtain preceding and following temporal image slice pairs with visual guidance markers; Construct system prompt words based on the interpretation rules for natural resource monitoring operations; Construct inference instructions based on the attachment attributes of the target patch data; The system prompts, inference instructions, and preceding and following time-phase image slices are input to a pre-adapted, multimodal large model. Multimodal fusion inference processing is then performed to obtain the authenticity judgment results of the changed patches and the corresponding inference logic chain. Based on the authenticity judgment results, false change patch removal processing is performed to complete the false alarm suppression of natural resource monitoring.

[0007] In one possible implementation, the step of performing spatial overlay analysis on the changed patch vector to obtain the target patch data with attached business attributes includes: The target layer is the vector of changed land parcels, and the land use parcel layer and thematic element layer from the land change survey are the reference layers. Spatial overlay calculations are performed on the target layer and the reference layer to obtain the intersection area and proportion data between the layers; Based on the intersecting area and proportion data, the land type attributes and thematic element attributes of the reference layer are linked to the target layer to obtain the target patch data with the linked business attributes.

[0008] In one possible implementation, the step of performing spatial matching and cropping processing on the target patch data and the preceding and following temporal remote sensing images to obtain preceding and following temporal image slice pairs with visually guided markers includes: Spatial coordinate transformation is performed on the target patch data and the preceding and following temporal remote sensing images to obtain patch and image data in coordinate system one; The vector boundaries of the target patch data are mapped to the pixel space of the preceding and following temporal remote sensing images to calculate the dynamic cropping window; The dynamic cropping window is used to perform cropping processing on the preceding and following temporal remote sensing images to obtain an initial image slice that encapsulates the target patch data. Visual guidance mark overlay processing is performed on the initial image slices to obtain preceding and following temporal image slice pairs with visual guidance marks.

[0009] In one possible implementation, the step of constructing system prompt words based on natural resource monitoring operational interpretation rules includes: Review the interpretation rules and clauses in the Natural Resources Monitoring Operation Guidelines; The judgment rule clauses are subjected to generalization transformation processing to obtain the logical boundary instructions for model reasoning; By combining the reasoning character profile, task requirements, judgment rules, and output format requirements, system prompt words are generated.

[0010] In one possible implementation, the step of constructing inference instructions based on the attachment attributes of the target patch data includes: Obtain the land category attributes and thematic element attributes corresponding to each map patch in the target map patch data; Based on the preset instruction template, the land category attributes and thematic element attributes are filled into the instruction template to construct a personalized reasoning instruction that is bound to the unique identifier of each map patch.

[0011] In one possible implementation, the pre-tuned multimodal large model with domain adaptation is obtained through the following steps: Acquire historical verification patch data and corresponding remote sensing images, perform spatial overlay analysis and image cropping processing to obtain sample image slices and sample patch data with attached business attributes; Based on the business interpretation rules and the attached business attributes of the sample patch data, sample generation prompt words are constructed. Combined with the sample image slices, the pre-trained teacher model is called to perform inference processing to obtain an initial sample set with inference path. Logical correction and semantic refinement are performed on the initial sample set to obtain the final training sample set for image, attribute and logical chain matching; The final training sample set is used to perform efficient parameter fine-tuning on the general multimodal large model to obtain a domain-adapted multimodal large model.

[0012] In one possible implementation, the step of performing efficient parameter fine-tuning on a general multimodal large model using the final training sample set includes: A training dataset for the model is constructed by combining sample image slice paths, system prompts, inference instructions, and corrected inference text. Set the training validation ratio, multimodal parameter freezing strategy, number of training rounds, and fine-tuning method parameters; Based on the training dataset and the set parameters, supervised fine-tuning is performed on the general multimodal large model to obtain a domain-adapted multimodal large model.

[0013] In one possible implementation, the step of performing multimodal fusion inference processing on a pre-adapted multimodal large model by combining the system prompts, inference instructions, and preceding and following temporal image slices includes: A master-slave architecture is used to deploy a multimodal large model that is adapted to the domain, and a task library is built based on a spatial database; Write the system prompts, inference instructions and preceding and following time-phase image slices into the task library, and generate inference tasks in fixed batches. The inference task is distributed to multiple computing nodes through a message queue to perform parallel multimodal fusion inference processing, thereby obtaining the authenticity judgment results of each patch and the corresponding inference logic chain.

[0014] In one possible implementation, the step of performing pseudo-change patch removal processing based on the authenticity discrimination result includes: Based on the authenticity determination results, marking is performed on the patches that are determined to be false changes; Remove the patches marked as pseudo-changes and retain the patch data determined as true changes; The inference logic chain is associated with and stored with the map data of the actual changes to generate natural resource monitoring results data.

[0015] In one possible implementation, the step of performing visually guided marker overlay processing on the initial image slice includes: The boundary edge tensor of the target patch data is extracted using the geometric mask dilation algorithm; Pixel enhancement processing is performed on the boundary edge tensor to obtain a high-contrast boundary marker; The high-contrast boundary markers are superimposed onto the initial image slices to obtain a pair of front and back temporal image slices with visual guidance markers.

[0016] Secondly, the present invention provides a false alarm suppression device for natural resource monitoring based on a multimodal large model, comprising: The data acquisition unit is used to acquire change vectors of natural resource monitoring patches and corresponding preceding and following temporal remote sensing images; A spatial overlay analysis unit is used to perform spatial overlay analysis on the changed patch vector to obtain target patch data with attached business attributes; The matching and cropping unit is used to perform spatial matching and cropping processing on the target patch data and the preceding and following temporal remote sensing images to obtain preceding and following temporal image slice pairs with visual guidance marks. The prompt word construction unit is used to construct system prompt words according to the natural resource monitoring business interpretation rules and to construct reasoning instructions according to the attachment attributes in the target patch data. The fusion reasoning unit is used to input a multimodal large model that has been pre-adapted and fine-tuned for domain adaptation by combining the system prompts, reasoning instructions and previous and subsequent time-phase image slices, and perform multimodal fusion reasoning processing to obtain the authenticity judgment result of the changed patches and the corresponding reasoning logic chain. The monitoring unit is used to perform false change patch removal processing based on the true / false discrimination results, thereby completing the false alarm suppression of natural resource monitoring.

[0017] Compared with existing technologies, the advantages of this invention are as follows: This invention obtains operational attribute data by performing spatial overlay analysis on the vectors of changed patches, and combines this with the cropping results of preceding and following temporal remote sensing images, using both as input data for a multimodal large model. This allows the model inference process to simultaneously incorporate image visual features and semantic information from natural resource monitoring operations. Unlike existing technologies that rely solely on image spectral and texture features for false alarm suppression, this invention can distinguish between ground feature change scenarios with similar visual features but different operational attributes, reducing misjudgments of false change patches and improving the completeness of false alarm suppression.

[0018] This invention constructs system prompts by organizing the interpretation rules of natural resource monitoring operations, setting clear logical boundaries for the reasoning process of a multimodal large model, and guiding the model to output discrimination results that include the complete reasoning process. Unlike the end-to-end black-box discrimination models in existing technologies, the discrimination results output by this invention include interpretation criteria corresponding to image features and business rules, making the entire discrimination process traceable and directly supporting business review processes. Simultaneously, this invention transforms non-image-related business vector data into text instructions that the model can recognize, achieving the fusion of multi-source heterogeneous data at the reasoning level. Unlike existing technologies that only perform post-processing spatial overlay filtering on multi-source data, this invention can fully utilize various prior business knowledge to participate in the discrimination process, improving the reliability of the discrimination results.

[0019] This invention employs a master-slave architecture to deploy a domain-adapted multimodal large model. It utilizes message queues to distribute and parallelize inference tasks, enabling the synchronous identification of large batches of changing map features on conventional servers without relying on high-performance dedicated hardware. Unlike existing solutions that demand high-performance hardware, this invention reduces deployment costs while improving the processing efficiency of large-scale natural resource monitoring tasks, better meeting the needs of large-scale business scenarios. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a flowchart illustrating the overall process of the false alarm suppression method for natural resource monitoring based on a multimodal large model, as described in this invention. Figure 2 This is a schematic diagram of dynamic cropping of remote sensing images based on changing patch boundaries according to an embodiment of the present invention; Figure 3 This is a schematic diagram of visual guidance markings for the boundary images of changing patches according to an embodiment of the present invention; Figure 4 This is a schematic diagram of manual verification of the sample set and the human-computer interaction interface in an embodiment of the present invention; Figure 5 This is a schematic diagram of the multimodal large model distributed cluster inference architecture according to an embodiment of the present invention; Figure 6 This is a flowchart illustrating the false alarm suppression method for natural resource monitoring based on a multimodal large model, according to an embodiment of the present invention. Detailed Implementation

[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0023] Example:

[0024] It should be noted that the terms "comprising" and "having" and any variations thereof in the embodiments of the present invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such processes, methods, products, or devices.

[0025] See Figure 6 This embodiment provides a method for suppressing false alarms in natural resource monitoring based on a multimodal large model, including the following steps: Step 101: Obtain the vector of the change patch in the natural resource monitoring and the corresponding before and after time-phase remote sensing images.

[0026] Specifically, the change patch vector can be the vector data of the change area of ​​the land surface extracted by natural resource remote sensing monitoring; the previous and next time phase remote sensing images can be satellite remote sensing images of the same area acquired at different times, such as Gaofen-2 satellite images with a 2-year interval.

[0027] Step 102: Perform spatial overlay analysis on the changed patch vector to obtain the target patch data with attached business attributes.

[0028] Specifically, spatial overlay analysis can be a geographic calculation analysis of the spatial intersection relationship of vector layers; attaching business attributes can be associating business fields such as land type and thematic elements with changed map patches; target map patch data can be vector data of changed map patches with attached business attributes.

[0029] The steps of performing spatial overlay analysis on the changed patch vectors to obtain the target patch data for attaching business attributes include: The target layer is the vector of changed land parcels, and the land use parcel layer and thematic element layer from the land change survey are the reference layers. Spatial overlay calculations are performed on the target layer and the reference layer to obtain the intersection area and proportion data between the layers; Based on the intersecting area and proportion data, the land type attributes and thematic element attributes of the reference layer are linked to the target layer to obtain the target patch data with the linked business attributes.

[0030] Specifically, the change patch vector can be a change area vector layer formed by natural resource monitoring; the land use patch layer from the land change survey can be an annual land use status classification vector layer; the thematic element layer can be a vector layer from special monitoring such as fill areas, photovoltaic areas, and areas with unfinished demolition; spatial overlay calculation processing can be a geographic analysis operation that calculates the intersection area and proportion of vector layers; the intersection area and proportion data can be the overlapping area and proportion values ​​of the target patch and the reference patch; the land use attribute can be a land use classification field such as paddy field, dry land, orchard; the thematic element attribute can be a special element field such as fill and photovoltaic facilities. For example, when the intersection area proportion is not less than 0.8, the corresponding thematic element attribute is attached; for patches that do not intersect with any thematic layer, the attribute field is assigned a null value.

[0031] Step 103: Perform spatial matching and cropping processing on the target patch data and the preceding and following temporal remote sensing images to obtain preceding and following temporal image slice pairs with visual guidance markers.

[0032] Specifically, spatial matching and cropping can be an operation that unifies the coordinates of the patch and the image and crops the local image; visual guidance markers can be high-contrast boundary markers superimposed on the image; and preceding and following temporal image slice pairs can be combinations of image slices of the same range in preceding and following time periods.

[0033] The step of performing spatial matching and cropping processing on the target patch data and the preceding and following temporal remote sensing images to obtain preceding and following temporal image slice pairs with visual guidance markers includes: Spatial coordinate transformation is performed on the target patch data and the preceding and following temporal remote sensing images to obtain patch and image data in coordinate system one; The vector boundaries of the target patch data are mapped to the pixel space of the preceding and following temporal remote sensing images to calculate the dynamic cropping window; The dynamic cropping window is used to perform cropping processing on the preceding and following temporal remote sensing images to obtain an initial image slice that encapsulates the target patch data. Visual guidance mark overlay processing is performed on the initial image slices to obtain preceding and following temporal image slice pairs with visual guidance marks.

[0034] Specifically, spatial coordinate transformation processing can be an operation that unifies the patch and the image to the 2000 National Geodetic Coordinate System; pixel space can be the pixel coordinate system of the remote sensing image; dynamic cropping window can be the cropping range calculated based on the bounding rectangle of the patch and the scale factor; initial image slice can be the original local image after cropping; visual guidance mark overlay processing can be an operation that overlays high-contrast boundary marks on the image; the scale factor of the dynamic cropping window can be 1.5; visual guidance marks can be patch boundary tensor marks enhanced by the red channel.

[0035] Further, the step of performing visually guided marker overlay processing on the initial image slice includes: The boundary edge tensor of the target patch data is extracted using the geometric mask dilation algorithm; Pixel enhancement processing is performed on the boundary edge tensor to obtain a high-contrast boundary marker; The high-contrast boundary markers are superimposed onto the initial image slices to obtain a pair of front and back temporal image slices with visual guidance markers.

[0036] Specifically, the geometric mask dilation algorithm can be an image processing algorithm for extracting the boundary contours of patches; the boundary edge tensor can be a pixel matrix representing the boundary contours of patches; pixel enhancement processing can be an image processing operation that improves the contrast of boundary pixels; and high-contrast boundary markers can be color-highlighted and easily identifiable patch boundary markers. For example, if pixel enhancement processing uses a red channel enhancement method, and high-contrast boundary markers are red outlines, the patch range and background area can be clearly distinguished after overlay.

[0037] Step 104: Construct system prompt words based on the natural resources monitoring business interpretation rules.

[0038] Specifically, system prompts can be textual instructions that define the rules for model inference.

[0039] The steps for constructing system prompt words based on the natural resource monitoring business interpretation rules include: Review the interpretation rules and clauses in the Natural Resources Monitoring Operation Guidelines; The judgment rule clauses are subjected to generalization transformation processing to obtain the logical boundary instructions for model reasoning; By combining the reasoning character profile, task requirements, judgment rules, and output format requirements, system prompt words are generated.

[0040] Specifically, the interpretation rules can be operational guidelines for distinguishing between genuine and pseudo-changes in natural resource monitoring; the generalization and transformation process can be the extraction of professional rules into instruction text that the model can recognize; the logical boundary instructions can be textual content that limits the scope of model reasoning and judgment criteria; the reasoning persona can be the role positioning of a senior natural resource internal verification officer; the task requirements can be the work requirements for completing the distinction between genuine and pseudo-changes based on imagery and attributes; and the output format requirements can be structured output requirements that include interpretation conclusions, interpretation reasons, and interpretation land types. For example, pseudo-change rules include situations such as seasonal decline and agricultural tillage, while genuine change rules include situations such as new construction and hardened ground.

[0041] Step 105: Construct inference instructions based on the attachment attributes in the target patch data.

[0042] Specifically, inference instructions can be personalized text instructions that combine the attributes of target patch data to guide multimodal large models to carry out interpretation and inference.

[0043] The step of constructing inference instructions based on the attachment attributes of the target patch data includes: Obtain the land category attributes and thematic element attributes corresponding to each map patch in the target map patch data; Based on the preset instruction template, the land category attributes and thematic element attributes are filled into the instruction template to construct a personalized reasoning instruction that is bound to the unique identifier of each map patch.

[0044] Specifically, the preset instruction template can be a standardized text template designed in advance based on the business needs of natural resource monitoring. The template contains fixed fill positions and guiding statements, which are used to fill in personalized information such as land type attributes and thematic element attributes. The guiding statements are used to clarify the core requirement that the model needs to conduct reasoning in conjunction with the attributes of the map patch.

[0045] Step 106: Input the system prompts, inference instructions, and preceding and following time-phase image slices into a pre-adapted multimodal large model, perform multimodal fusion inference processing, and obtain the authenticity judgment results of the changed patches and the corresponding inference logic chain.

[0046] Specifically, the multimodal large model can be a visual language large model that has been fine-tuned in the field of natural resources; the multimodal fusion reasoning processing can be semantic logic reasoning that combines image and text attributes; the true / false judgment result can be a judgment conclusion on whether the change is real or false; and the reasoning logic chain can be a text description containing the basis for judgment.

[0047] The multimodal large model that has undergone pre-adaptation and fine-tuning for domain adaptation is obtained through the following steps: Acquire historical verification patch data and corresponding remote sensing images, perform spatial overlay analysis and image cropping processing to obtain sample image slices and sample patch data with attached business attributes; Based on the business interpretation rules and the attached business attributes of the sample patch data, sample generation prompt words are constructed. Combined with the sample image slices, the pre-trained teacher model is called to perform inference processing to obtain an initial sample set with inference path. Logical correction and semantic refinement are performed on the initial sample set to obtain the final training sample set for image, attribute and logical chain matching; The final training sample set is used to perform efficient parameter fine-tuning on the general multimodal large model to obtain a domain-adapted multimodal large model.

[0048] Specifically, historical verification map data can be vector data of changed map patches from previous natural resource monitoring that have been manually verified; the pre-trained teacher model can be a large-scale visual language pre-trained model that supports internet interface calls or local deployment; the initial sample set can be a sample set containing images, attributes, inference conclusions, and inference paths; logical correction and semantic polishing can be operations that manually correct model inference illusions and standardize inference text; efficient parameter fine-tuning can be a model fine-tuning method with low resource consumption. For example, the pre-trained teacher model can be a general multimodal visual language large model, and the final training sample set can be a structured sample set containing image tiles, land use attributes, thematic elements, and standardized inference logic.

[0049] Furthermore, the steps for performing efficient parameter fine-tuning on the general multimodal large model using the final training sample set include: A training dataset for the model is constructed by combining sample image slice paths, system prompts, inference instructions, and corrected inference text. Set the training validation ratio, multimodal parameter freezing strategy, number of training rounds, and fine-tuning method parameters; Based on the training dataset and the set parameters, supervised fine-tuning is performed on the general multimodal large model to obtain a domain-adapted multimodal large model.

[0050] Specifically, the model training dataset can be a structured data file containing sample paths, text instructions, and inference labels; the training-to-validation ratio can be the ratio of training samples to test samples; the multimodal parameter freezing strategy can be freezing the parameter settings of the visual encoder, multimodal projector, or language model; the number of training epochs can be the number of times the model traverses all training samples; and the fine-tuning method parameters can be relevant parameters of low-rank adaptation. For example, the model training dataset can be stored in JSON format, the training-to-validation ratio can be set to 8:2, the number of training epochs can be set to 1 to 3, and the fine-tuning method can be low-rank adaptation.

[0051] The steps of performing multimodal fusion inference processing on the multimodal large model that has been pre-adapted and fine-tuned for domain adaptation by combining the system prompts, inference instructions, and preceding and following time-phase image slices include: A master-slave architecture is used to deploy a multimodal large model that is adapted to the domain, and a task library is built based on a spatial database; Write the system prompts, inference instructions and preceding and following time-phase image slices into the task library, and generate inference tasks in fixed batches. The inference task is distributed to multiple computing nodes through a message queue to perform parallel multimodal fusion inference processing, thereby obtaining the authenticity judgment results of each patch and the corresponding inference logic chain.

[0052] Specifically, the master-slave architecture can be a distributed deployment structure where the master node manages tasks and the slave nodes execute inference; the spatial database can be a database that supports spatial data storage and querying; the task repository can be a data table that stores inference tasks to be processed; the message queue can be middleware that enables asynchronous task distribution; and the compute nodes can be inference servers equipped with graphics processors. For example, the spatial database can use a PostGIS database, the message queue can use a Celery message queue, the inference deployment can use a vLLM framework, and the compute nodes can be multiple graphics processor servers.

[0053] Step 107: Perform false change patch removal processing based on the true / false discrimination results to complete the false alarm suppression of natural resource monitoring.

[0054] Specifically, the false change patch removal process can be the operation of removing false alarm patches from monitoring results.

[0055] The steps for removing false change patches based on the authenticity determination results include: Based on the authenticity determination results, marking is performed on the patches that are determined to be false changes; Remove the patches marked as pseudo-changes and retain the patch data determined as true changes; The inference logic chain is associated with and stored with the map data of the actual changes to generate natural resource monitoring results data.

[0056] Specifically, the labeling process can be the operation of adding pseudo-change attribute identifiers to map features; the removal operation can be the operation of removing pseudo-change map features from the monitoring dataset; the associated storage process can be the operation of binding and storing the inference logic chain with the spatial data of map features; and the natural resource monitoring results data can be standardized monitoring results containing true change map features and the inference basis. For example, pseudo-change map feature labeling uses dedicated field identifiers, the removal operation is implemented through spatial data filtering, and the associated storage uses a relational database to complete the field binding.

[0057] The specific embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. The false alarm suppression method for natural resource monitoring based on a multimodal large model proposed in this embodiment consists of three parts: sample dataset construction, multimodal large model fine-tuning, and distributed cluster inference. The overall process route is as follows: Figure 1 exhibit.

[0058] The sample dataset is constructed using historical verification map patches and reference data as initial data. First, a business attribute attachment operation based on spatial overlay is performed. The historically identified change map patch vector is used as the target layer, and the annual land use change survey land use map patch layer and thematic element layer are used as reference layers. The intersection area and ratio between the target layer and the reference layer are calculated through overlay analysis. According to the principle of maximizing the area ratio, the land use attributes of the land use map patch layer are attached to the target layer. At the same time, the thematic element attributes of the intersection area that meets the set conditions are attached to the target layer. The corresponding attributes of the layer elements that do not intersect with the thematic layer are assigned null values.

[0059] After completing the business attribute attachment, an adaptive image cropping operation is performed. Using the changed patch and its corresponding preceding and following temporal remote sensing images as the data source, a spatial coordinate transformation is performed on the changed patch and the preceding and following temporal remote sensing images to place the patch and the image in the same coordinate system. Then, the patch vector boundary is mapped to the pixel space of the preceding and following temporal remote sensing images. Using the minimum bounding rectangle of the patch as a reference, and combining a 1.5x scale factor, a dynamic cropping window is calculated. The dynamic cropping window can completely encompass the target feature while preserving the surrounding geographic context. The image dynamic cropping process based on the changed patch boundary is achieved through… Figure 2 exhibit.

[0060] After image cropping, the geometric mask dilation algorithm is used to extract the edge tensor of the patch boundaries. High-contrast pixels are then superimposed onto the image slices to generate enhanced before-and-after temporal image slice pairs with visual guidance markers. The visual guidance markers for the changing patch boundaries are then... Figure 3The demonstration then proceeds. Based on the natural resource monitoring operation guidelines or expert experience, the interpretation rules are summarized and transformed into generalized system prompts. These system prompts include the reasoning persona, task content, interpretation rules, and output format, which are used to establish the logical boundaries of the model's reasoning.

[0061] After constructing the system prompts, a sample set generation operation is performed. Logical chains are constructed using the attached business attributes to generate guiding instructions. These instructions are then filled into a pre-defined initial sample set to generate a guiding instruction template. Simultaneously, historical judgment records are inserted into the template as prior knowledge. The system prompts and image slices are input into a large-scale pre-trained teacher model to generate inference text with logical chains. A full check and random sampling review of the initial sample set are performed through human-computer interaction to correct inference illusions and conclusion biases generated by the model, resulting in high-confidence samples. The manual verification process for the sample set is implemented through a human-computer interactive verification interface. Figure 4 exhibit.

[0062] After the sample dataset is constructed, a multimodal large model domain adaptation and fine-tuning operation is performed. A general multimodal large model that supports dynamic resolution input is selected as the base. A set number and proportion of true and false change samples are selected for model training. A portion of the samples are reserved as a test set for result evaluation. A dataset file is constructed by combining image slice paths, system prompts, business application inference instructions, and expert-corrected inference text. Parameters such as the validation set ratio, multimodal parameter freezing strategy, number of training rounds, and fine-tuning method are set. Supervised fine-tuning training is performed using parameter-efficient fine-tuning techniques to complete the domain adaptation of the general multimodal large model and obtain an inference model with expert interpretation capabilities. At the same time, the model performance is evaluated, business misjudgment feedback is received, and the model effect is optimized.

[0063] After model fine-tuning, distributed cluster inference is performed. A domain-adapted inference model is deployed using a master-slave architecture. A task library is built based on a spatial database, and the maps to be inferred and reference data are written into the task library. Tasks are read and distributed in fixed batch order. A message queue enables computing nodes to compete for tasks, achieving parallel processing and high-throughput interpretation. The inference process is completed using an inference framework. After the maps to be judged complete preliminary processing, the distributed cluster is invoked to perform multimodal fusion inference based on system prompts, outputting true / false judgment results with logical chains. After inference, the structured results are written to the database. The distributed cluster inference architecture of the multimodal large model is implemented through... Figure 5 exhibit.

[0064] This embodiment completes the sample dataset construction, multimodal large model fine-tuning, and distributed cluster inference sequentially through the above steps. Finally, false change patches are removed based on the true / false discrimination results to achieve false alarm suppression in natural resource monitoring. At the same time, the inference logic chain is associated with and stored with the true change patch data to form natural resource monitoring results data.

[0065] Based on the same inventive concept, this embodiment also provides a false alarm suppression device for natural resource monitoring based on a multimodal large model, comprising: The data acquisition unit is used to acquire change vectors of natural resource monitoring patches and corresponding preceding and following temporal remote sensing images; A spatial overlay analysis unit is used to perform spatial overlay analysis on the changed patch vector to obtain target patch data with attached business attributes; The matching and cropping unit is used to perform spatial matching and cropping processing on the target patch data and the preceding and following temporal remote sensing images to obtain preceding and following temporal image slice pairs with visual guidance marks. The prompt word construction unit is used to construct system prompt words according to the natural resource monitoring business interpretation rules and to construct reasoning instructions according to the attachment attributes in the target patch data. The fusion reasoning unit is used to input a multimodal large model that has been pre-adapted and fine-tuned for domain adaptation by combining the system prompts, reasoning instructions and previous and subsequent time-phase image slices, and perform multimodal fusion reasoning processing to obtain the authenticity judgment result of the changed patches and the corresponding reasoning logic chain. The monitoring unit is used to perform false change patch removal processing based on the true / false discrimination results, thereby completing the false alarm suppression of natural resource monitoring.

[0066] Since this device corresponds to the false alarm suppression method for natural resource monitoring based on a multimodal large model in this embodiment of the invention, and the principle of this device in solving the problem is similar to that of this method, the implementation of this device can refer to the implementation process of the above method embodiment, and the repeated parts will not be described again.

[0067] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0068] The above embodiments are merely illustrative of the technical concept and features of the present invention, and are intended to enable those skilled in the art to understand the content of the present invention and implement it accordingly. They should not be construed as limiting the scope of protection of the present invention. All equivalent changes or modifications made based on the essence of the content of the present invention should be covered within the scope of protection of the present invention.

Claims

1. A method for suppressing false alarms in natural resource monitoring based on a multimodal large model, characterized in that, Includes the following steps: Acquire vectors of changes in natural resource monitoring patches and their corresponding preceding and following temporal remote sensing images; Perform spatial overlay analysis on the changed patch vectors to obtain target patch data for the attached service attributes; Spatial matching and cropping processing is performed on the target patch data and the preceding and following temporal remote sensing images to obtain preceding and following temporal image slice pairs with visual guidance markers; Construct system prompt words based on the interpretation rules for natural resource monitoring operations; Construct inference instructions based on the attachment attributes of the target patch data; The system prompts, inference instructions, and preceding and following time-phase image slices are used to input a pre-adapted, multimodal large model. Multimodal fusion inference processing is then performed to obtain the authenticity judgment results of the changed patches and the corresponding inference logic chain. Based on the authenticity judgment results, false change patch removal processing is performed to complete the false alarm suppression of natural resource monitoring.

2. The method according to claim 1, characterized in that, The steps for performing spatial overlay analysis on the changed patch vectors to obtain the target patch data with attached business attributes include: The target layer is the vector of changed land parcels, and the land use parcel layer and thematic element layer from the land change survey are the reference layers. Spatial overlay calculations are performed on the target layer and the reference layer to obtain the intersection area and proportion data between the layers; Based on the intersecting area and proportion data, the land type attributes and thematic element attributes of the reference layer are linked to the target layer to obtain the target patch data with the linked business attributes.

3. The method according to claim 1, characterized in that, The steps of performing spatial matching and cropping processing on the target patch data and the preceding and following temporal remote sensing images to obtain preceding and following temporal image slice pairs with visual guidance markers include: Spatial coordinate transformation is performed on the target patch data and the preceding and following temporal remote sensing images to obtain patch and image data in coordinate system one; The vector boundaries of the target patch data are mapped to the pixel space of the preceding and following temporal remote sensing images to calculate the dynamic cropping window; The dynamic cropping window is used to perform cropping processing on the preceding and following temporal remote sensing images to obtain an initial image slice that encapsulates the target patch data. Visual guidance mark overlay processing is performed on the initial image slices to obtain preceding and following temporal image slice pairs with visual guidance marks.

4. The method according to claim 1, characterized in that, The steps for constructing system prompt words based on the natural resource monitoring operational interpretation rules include: Review the interpretation rules and clauses in the Natural Resources Monitoring Operation Guidelines; The judgment rule clauses are subjected to generalization transformation processing to obtain the logical boundary instructions for model reasoning; By combining the reasoning character profile, task requirements, judgment rules, and output format requirements, system prompts are generated. The steps for constructing inference instructions based on the attachment attributes of the target patch data include: Obtain the land category attributes and thematic element attributes corresponding to each map patch in the target map patch data; Based on the preset instruction template, the land category attributes and thematic element attributes are filled into the instruction template to construct a personalized reasoning instruction that is bound to the unique identifier of each map patch.

5. The method according to claim 1, characterized in that, The pre-adapted, fine-tuned multimodal large model is obtained through the following steps: Acquire historical verification patch data and corresponding remote sensing images, perform spatial overlay analysis and image cropping processing to obtain sample image slices and sample patch data with attached business attributes; Based on the business interpretation rules and the attached business attributes of the sample patch data, sample generation prompt words are constructed. Combined with the sample image slices, the pre-trained teacher model is called to perform inference processing to obtain an initial sample set with inference path. Logical correction and semantic refinement are performed on the initial sample set to obtain the final training sample set for image, attribute and logical chain matching; The final training sample set is used to perform efficient parameter fine-tuning on the general multimodal large model to obtain a domain-adapted multimodal large model.

6. The method according to claim 5, characterized in that, The steps for efficiently fine-tuning the parameters of a general multimodal large model using the final training sample set include: A training dataset for the model is constructed by combining sample image slice paths, system prompts, inference instructions, and corrected inference text. Set the training validation ratio, multimodal parameter freezing strategy, number of training rounds, and fine-tuning method parameters; Based on the training dataset and the set parameters, supervised fine-tuning is performed on the general multimodal large model to obtain a domain-adapted multimodal large model.

7. The method according to claim 1, characterized in that, The steps of performing multimodal fusion inference processing on the pre-adapted multimodal large model, which is input with the system prompts, inference instructions, and preceding and following time-phase image slices, include: A master-slave architecture is used to deploy a multimodal large model that is adapted to the domain, and a task library is built based on a spatial database; Write the system prompts, inference instructions and preceding and following time-phase image slices into the task library, and generate inference tasks in fixed batches. The inference task is distributed to multiple computing nodes through a message queue to perform parallel multimodal fusion inference processing, thereby obtaining the authenticity judgment results of each patch and the corresponding inference logic chain.

8. The method according to claim 1, characterized in that, The steps for removing pseudo-change patches based on the authenticity determination results include: Based on the authenticity determination results, marking is performed on the patches that are determined to be false changes; Remove the patches marked as pseudo-changes and retain the patch data determined as true changes; The inference logic chain is associated with and stored with the map data of the actual changes to generate natural resource monitoring results data.

9. The method according to claim 3, characterized in that, The steps of performing visually guided mark overlay processing on the initial image slices include: The boundary edge tensor of the target patch data is extracted using the geometric mask dilation algorithm; Pixel enhancement processing is performed on the boundary edge tensor to obtain a high-contrast boundary marker; The high-contrast boundary markers are superimposed onto the initial image slices to obtain a pair of front and back temporal image slices with visual guidance markers.

10. A false alarm suppression device for natural resource monitoring based on a multimodal large model, characterized in that, include: The data acquisition unit is used to acquire change vectors of natural resource monitoring patches and corresponding preceding and following temporal remote sensing images; A spatial overlay analysis unit is used to perform spatial overlay analysis on the changed patch vector to obtain target patch data with attached business attributes; The matching and cropping unit is used to perform spatial matching and cropping processing on the target patch data and the preceding and following temporal remote sensing images to obtain preceding and following temporal image slice pairs with visual guidance marks. The prompt word construction unit is used to construct system prompt words according to the natural resource monitoring business interpretation rules and to construct reasoning instructions according to the attachment attributes in the target patch data. The fusion reasoning unit is used to input a multimodal large model that has been pre-adapted and fine-tuned for domain adaptation by combining the system prompts, reasoning instructions and previous and subsequent time-phase image slices, and perform multimodal fusion reasoning processing to obtain the authenticity judgment result of the changed patches and the corresponding reasoning logic chain. The monitoring unit is used to perform false change patch removal processing based on the true / false discrimination results, thereby completing the false alarm suppression of natural resource monitoring.