Remote sensing task optimization method, model training method and device, equipment, medium and product
By constructing a training dataset containing multimodal remote sensing data and text task instructions, and training a large multimodal model, the problem of insufficient complex task decomposition and planning capabilities in existing technologies is solved, and the efficient execution of multi-step remote sensing tasks and the improvement of task generalization capabilities are achieved.
Patent Information
- Application Number
- CN202511589041.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-31
- Publication Date
- 2026-02-13
AI Technical Summary
Existing multimodal large models lack the ability to decompose and plan complex tasks in remote sensing applications, cannot accurately deduce the execution steps corresponding to complex tasks, and are ill-suited for comprehensive tasks involving multi-step collaboration.
By constructing a training dataset containing multimodal remote sensing data and text task instructions, a large multimodal model is trained to understand the semantic relationship between remote sensing data of different modalities and text under complex text tasks, automatically inferring task execution steps, and improving its ability to analyze and reason about complex tasks by correcting model parameters.
This enables multimodal large models to decompose and plan complex task instructions in remote sensing missions, improving the model's flexibility and adaptability, ensuring that the inference results of task execution steps meet actual needs, and enhancing task generalization ability and execution accuracy.
Smart Images

Figure CN121524913A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of computer vision and artificial intelligence, including but not limited to remote sensing cognition and perception collaborative optimization methods based on multimodal large models, training methods, devices, equipment, media and products of multimodal large models. Background Technology
[0002] In recent years, with breakthroughs in multimodal large language model technology, intelligent analysis methods based on generative pre-training frameworks have brought revolutionary opportunities to the field of remote sensing. These models, by constructing a unified visual-language joint representation space, have initially achieved cross-modal semantic alignment and task generalization capabilities. However, current multimodal large models applied in remote sensing often employ an end-to-end single-round inference paradigm in practical applications, lacking the ability to decompose and plan complex tasks. They cannot accurately infer the execution steps corresponding to complex tasks, thus making them unsuitable for comprehensive tasks requiring multi-step collaboration. Summary of the Invention
[0003] This application provides a remote sensing cognition and perception collaborative optimization method based on a multimodal large model, a training method, apparatus, device, medium, and product for the multimodal large model; wherein, In a first aspect, embodiments of this application provide a remote sensing cognition and perception collaborative optimization method based on a multimodal large model. The method includes: based on a remote sensing image to be processed and a first text task instruction, obtaining the task execution step inference result corresponding to the first text task instruction through a multimodal large model; wherein, the multimodal large model is obtained by training a training dataset; the training dataset includes a remote sensing image dataset and a task instruction dataset; the remote sensing image dataset includes multimodal remote sensing data; the task instruction dataset includes a second text task instruction and the task execution steps corresponding to the second text task instruction; and performing remote sensing task processing based on the task execution step inference result.
[0004] Secondly, embodiments of this application provide a training method for a multimodal large model, the method comprising: obtaining a task execution step prediction result based on a training dataset and an initial multimodal large model; wherein the training dataset includes a remote sensing image dataset and a task instruction dataset; the remote sensing image dataset includes multimodal remote sensing data; the task instruction dataset includes a second text task instruction and task execution steps corresponding to the second text task instruction; and correcting the initial multimodal large model based on the task execution step prediction result and the task execution steps to obtain a multimodal large model; wherein the multimodal large model is used to obtain the task execution step inference result corresponding to the first text task instruction through the remote sensing image to be processed and the first text task instruction.
[0005] Thirdly, embodiments of this application provide a remote sensing cognition and perception collaborative optimization device based on a multimodal large model. The device includes: a processing unit configured to obtain the task execution step inference result corresponding to the first text task instruction through the multimodal large model based on the remote sensing image to be processed and the first text task instruction; wherein, the multimodal large model is obtained by training through a training dataset; the training dataset includes a remote sensing image dataset and a task instruction dataset; the remote sensing image dataset includes multimodal remote sensing data; the task instruction dataset includes a second text task instruction and the task execution steps corresponding to the second text task instruction; and an execution unit configured to perform remote sensing task processing based on the task execution step inference result.
[0006] Fourthly, embodiments of this application provide a training apparatus for a multimodal large model. The apparatus includes: a training unit configured to obtain task execution step prediction results based on a training dataset and an initial multimodal large model; wherein the training dataset includes a remote sensing image dataset and a task instruction dataset; the remote sensing image dataset includes multimodal remote sensing data; the task instruction dataset includes a second text task instruction and task execution steps corresponding to the second text task instruction; and a correction unit configured to correct the initial multimodal large model based on the task execution step prediction results and the task execution steps to obtain a multimodal large model; wherein the multimodal large model is used to obtain the task execution step inference results corresponding to the first text task instruction through the remote sensing image to be processed and the first text task instruction.
[0007] Fifthly, embodiments of this application provide an electronic device, including a memory and a processor, wherein the memory stores a computer program that can run on the processor, and the processor executes the program to implement the method of the first aspect or the second aspect.
[0008] In a sixth aspect, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method of the first aspect or the second aspect.
[0009] In a seventh aspect, embodiments of this application provide a computer program product, including a computer program or instructions, which, when executed by a processor, implement the method of the first or second aspect of this application.
[0010] In this embodiment, the remote sensing image to be processed and the first text task instruction are first input into the multimodal large model to obtain the task execution step inference result corresponding to the first text task instruction; then, the remote sensing task processing is performed according to the task execution step inference result. Since the multimodal large model is trained based on remote sensing data containing multiple modalities, the second text task instruction, and the task execution steps corresponding to the second text task instruction, the multimodal large model can understand the relationship between the input of remote sensing data of different modalities and the text semantics under complex text tasks. This ensures that when applying the multimodal large model, only remote sensing data and text task instructions need to be input, and the task execution step inference corresponding to the text task instruction can be automatically inferred, thereby accurately executing comprehensive tasks that require multi-step coordination.
[0011] In this embodiment, the initial multimodal large model is first trained based on the training dataset to obtain the task execution step prediction results. Then, the parameters of the initial multimodal large model are continuously corrected based on the task execution step prediction results and the task execution steps in the training dataset. This allows the initial multimodal large model to continuously improve its ability to analyze and reason about complex task instructions during the training process, so that the inferred task execution step prediction results are increasingly close to the actual task execution steps. This ensures that when executing task instructions in different scenarios, it is not necessary to set up a separate training dataset, effectively ensuring that the multimodal large model has stronger task generalization ability and higher execution accuracy.
[0012] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0013] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the technical solutions of this application. Obviously, the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0014] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0015] Figure 1 This is a schematic diagram of the implementation process of the collaborative optimization method proposed in the embodiments of this application. Figure 1 ; Figure 2This is a schematic diagram of the structure of the multimodal large model proposed in the embodiments of this application; Figure 3 This is a schematic diagram illustrating the implementation process of obtaining the reasoning results of the task execution steps as proposed in an embodiment of this application; Figure 4 This illustration shows the process of co-optimization of cognition and perception in the multimodal large model proposed in this application. Figure 1 ; Figure 5 This illustration shows the process of co-optimization of cognition and perception in the multimodal large model proposed in this application. Figure 2 ; Figure 6 This is a schematic diagram illustrating the implementation process of the training method proposed in the embodiments of this application; Figure 7 This is a schematic diagram of the implementation process of the collaborative optimization method proposed in the embodiments of this application. Figure 2 ; Figure 8 This is a schematic diagram of the collaborative optimization device proposed in the embodiments of this application; Figure 9 This is a schematic diagram of the training device proposed in the embodiments of this application; Figure 10 This is a schematic diagram of the structure of the electronic device proposed in the embodiments of this application. Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the specific technical solutions of this application will be further described in detail below with reference to the accompanying drawings of the embodiments of this application. The following embodiments are used to illustrate this application, but are not intended to limit the scope of this application.
[0017] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0018] In the following description, references to "some embodiments," "this embodiment," "this application embodiment," and examples, etc., describe a subset of all possible embodiments. However, it is understood that "some embodiments" may be the same subset or different subset of all possible embodiments and may be combined with each other without conflict.
[0019] The descriptions such as "first," "second," and "third" appearing in the embodiments of this application do not have a specific meaning (such as no order, nor do they indicate a special limitation on the number of devices in the embodiments of this application), but are merely for the purpose of clearly describing the embodiments of this application and do not constitute any limitation on the embodiments of this application.
[0020] To facilitate understanding of the technical solutions of the embodiments of this application, the relevant technologies or terms of the embodiments of this application are described below. The following relevant technologies or terms are optional solutions and can be combined with the technical solutions of the embodiments of this application in any way, and all of them fall within the protection scope of the embodiments of this application.
[0021] In related technologies, traditional remote sensing intelligent analysis technology is mainly based on discriminative models designed specifically for single tasks, such as target detection frameworks or semantic segmentation algorithms based on deep convolutional neural networks.
[0022] Then, further research and analysis of the aforementioned technologies revealed that while these models achieve good performance in specific scenarios and pre-defined tasks, their inherent limitations become increasingly apparent as application scenarios become more complex and task requirements diversify. First, these models typically adopt a "one task, one model" design philosophy, strictly limiting the model architecture and parameter space to the specific task scope, making them unable to flexibly adapt to the complex and dynamically changing task requirements in real-world applications. Second, when it is necessary to expand to new task types or target categories, it is often necessary to redesign the model architecture or perform full-scale parameter fine-tuning, which not only leads to huge consumption of computing resources but also causes the model deployment and maintenance costs to increase exponentially.
[0023] In related technologies, traditional initial multimodal large models only design corresponding training datasets for a single task during the training process to complete the training of the initial multimodal large model.
[0024] Further research and analysis of the aforementioned technologies revealed that while these models achieve good performance in specific scenarios and for simple, pre-defined tasks, as application scenarios become more complex and task requirements diversify, directly incorporating complex task instructions during training results in inference results that deviate significantly from the desired outcomes. To achieve the desired training results, it becomes necessary to redesign the training dataset, leading to a massive workload and wasted resources.
[0025] In view of this, this application proposes a remote sensing cognition and perception collaborative optimization method based on a multimodal large model and a training method for the multimodal large model. In the collaborative optimization method of this application, the input remote sensing image to be processed and the first text task instruction are processed by the multimodal large model to obtain the task execution step inference result corresponding to the first text task instruction. Then, the remote sensing task processing is performed based on the task execution step inference result. Since the multimodal large model uses not only multimodal remote sensing data and the second text task instruction during training, but also the task execution steps corresponding to the second text task instruction, the multimodal large model possesses the ability to decompose and plan complex task instructions in practical applications, thereby ensuring that the output task execution step inference result meets the corresponding task requirements.
[0026] In the training method of this application, when constructing the training dataset of the initial multimodal large model, task execution steps corresponding to the second text task instructions are added to the training dataset containing multimodal remote sensing data and second text task instructions. This enables the initial multimodal large model to have the ability to plan and decompose complex tasks during the training process. Then, the parameters of the initial multimodal large model are continuously corrected based on the predicted results of the inferred task execution steps and the task execution steps. This makes the predicted results of the task execution steps inferred by the initial multimodal large model during the training process infinitely close to, or even the actual task execution steps, thus avoiding the need to design separate training datasets for different task scenarios.
[0027] One embodiment of this application provides a remote sensing cognition and perception collaborative optimization method based on a multimodal large model, which is applied to a remote sensing cognition and perception collaborative optimization device or electronic device based on a multimodal large model.
[0028] Figure 1 Schematic diagram of the implementation process of the collaborative optimization method provided in the embodiments of this application Figure 1 .like Figure 1 As shown, the method includes the following steps 101 to 102: Step 101: Based on the remote sensing image to be processed and the first text task instruction, obtain the task execution step inference result corresponding to the first text task instruction through a multimodal large model; wherein, the multimodal large model is obtained by training a training dataset; the training dataset includes a remote sensing image dataset and a task instruction dataset; the remote sensing image dataset includes multimodal remote sensing data; the task instruction dataset includes a second text task instruction and the task execution steps corresponding to the second text task instruction.
[0029] Step 102: Perform remote sensing task processing based on the reasoning results of the task execution steps.
[0030] It is understood that, in this embodiment of the application, by introducing a multimodal remote sensing dataset, a second text task instruction, and corresponding task execution steps from the training dataset, the multimodal large model, after training, is able to process multimodal remote sensing images and text information, while also possessing the ability to decompose and plan text information. Therefore, in actual implementation, the multimodal large model first receives the input remote sensing image to be processed and the first text task instruction, extracts and encodes features from the remote sensing image, and then, based on the content of the first text task instruction, infers multiple sub-task steps required to complete the first text task instruction, i.e., outputs the corresponding task execution step inference results. Thus, the corresponding remote sensing task is executed according to the task instruction step inference results. In this way, on the one hand, the multimodal large model gains text information segmentation and planning capabilities when handling complex tasks; on the other hand, since the inferred task execution steps are dynamically generated based on specific task instructions, the flexibility and adaptability of the multimodal large model are improved, enabling it to cope with task scenarios of varying complexity.
[0031] The following sections will describe further optional implementation methods for each of the above steps, as well as related terms.
[0032] In step 101, based on the remote sensing image to be processed and the first text task instruction, the task execution step inference result corresponding to the first text task instruction is obtained through a multimodal large model; wherein, the multimodal large model is obtained by training a training dataset; the training dataset includes a remote sensing image dataset and a task instruction dataset; the remote sensing image dataset includes multimodal remote sensing data; the task instruction dataset includes a second text task instruction and the task execution steps corresponding to the second text task instruction.
[0033] It should be understood that, in this embodiment, the remote sensing image to be processed is the remote sensing image input into the multimodal large model during actual application, and the multimodal large model performs corresponding remote sensing task processing based on this remote sensing image; the first text task instruction and the second text task instruction are instructions issued during actual application and training, respectively, based on the remote sensing task. The multimodal large model has at least the ability to recognize and understand corresponding content in the remote sensing image, and the ability to plan and understand the input task instructions. Based on this, the structure of the multimodal large model includes at least a visual encoder for extracting visual features from the remote sensing image, a visual language adapter for converting visual features into word vector features, and a large language model capable of processing text tasks. In the field of remote sensing, remote sensing tasks mainly include cognitive tasks and perceptual tasks, and may even simultaneously include both cognitive and perceptual tasks.
[0034] For example, a perception task could be to detect which targets are contained in a remote sensing image, what type of targets they are, and the spatial relationships between them; a cognitive task could be to decompose task instructions, the steps required to get from target point A in a remote sensing image to target point B in the same remote sensing image, and how to get from target point A to target point B; a remote sensing task that has both cognitive and perception tasks could be: how to get from target point C in one remote sensing image to target point D in another remote sensing image, to obtain exhibits in target point C and target point D respectively, and to classify these exhibits, etc.
[0035] In this embodiment of the application, the remote sensing image to be processed and the multimodal remote sensing data include images acquired through remote sensing technology that reflect the electromagnetic wave characteristics of targets on the Earth's surface or in the atmosphere. The remote sensing image to be processed is at least one mode of the multimodal remote sensing data.
[0036] For example, the remote sensing image to be processed includes optical images; multimodal remote sensing data includes optical images, radar images, and infrared images.
[0037] It should be understood that, in the embodiments of this application, remote sensing images include, but are not limited to, remote sensing images in the above three modes. Any image obtained through the corresponding remote sensing technology should be considered a remote sensing image of this application.
[0038] It should be understood that, in the embodiments of this application, the task instructions contained in the first text task instruction and the second text task instruction may or may not have a corresponding relationship with the remote sensing image.
[0039] For example, a remote sensing image includes a specific starting point, while the task instruction specifies the target point with specific coordinates and the waypoints to reach that target point; that is, the task instruction does not correspond to the remote sensing image. Alternatively, a remote sensing image may include a specific starting point and an ending point, while the task instruction specifies the waypoints in the remote sensing image to reach the ending point; that is, the task instruction corresponds to the remote sensing image.
[0040] In this embodiment, the task execution step inference result includes the task execution steps inferred by the multimodal large model to process the first text task instruction; the task execution steps corresponding to the second text task instruction include the actual task execution steps required to process the second text task instruction.
[0041] In step 102, remote sensing task processing is performed based on the task execution step inference result; wherein, the task execution step inference result of this application is infinitely close to, or even exactly, the actual task execution step. Therefore, the corresponding remote sensing task processing is performed according to the task execution step inference result.
[0042] For example, the first text task instruction includes how to get from target attraction C to target attraction D, and how to photograph exhibit U after arriving at target attraction D; the task execution step reasoning result includes going from target attraction C to waypoint M, then from waypoint M to target attraction D, and photographing exhibit U in target attraction D; the remote sensing task processing is to use an autonomous driving vehicle equipped with a high-definition camera to go from target attraction C to waypoint M, then from waypoint M to target attraction D, and finally use the vehicle-mounted high-definition camera to photograph exhibit U.
[0043] In some embodiments, Figure 2 A schematic diagram of the structure of a multimodal large model provided in this application embodiment is shown below. Figure 2 As shown, the multimodal large model 1000 includes a data padding module 1001 for adjusting the resolution of remote sensing images, an embedding module 1002 for dividing remote sensing images into image patches, a visual encoder 1003 for extracting visual features of remote sensing images, a visual language adapter 1004 for mapping visual features to word vector space, and a large language model 1005 for processing text tasks.
[0044] In some embodiments, the data dynamic filling module is a preprocessing module that automatically selects a suitable target resolution based on the actual resolution of the input remote sensing image and adjusts the actual resolution of the remote sensing image to the target resolution. Through the automatic resolution adjustment of the data dynamic filling module, it can effectively adapt to the input of remote sensing images with different aspect ratios, avoiding information loss or distortion caused by fixed-resolution cropping, and helping to improve the representation accuracy of remote sensing images.
[0045] In some embodiments, dividing the resolution-adjusted remote sensing image into image blocks by embedding a module can increase the processing efficiency of the visual encoder and enhance its ability to capture details.
[0046] In some embodiments, the visual encoder is designed based on the Transformer architecture, possessing powerful nonlinear modeling capabilities and cross-modal perception capabilities, and is responsible for efficiently extracting and representing image features from image patches in remote sensing images.
[0047] In some embodiments, when the visual language adapter processes image features, it progressively compresses and maps high-dimensional image features to word vector space through parametric projection, while retaining key semantic information of remote sensing images. This ensures that image content can be understood and utilized in the large language model, providing basic support for the language reasoning capabilities of the subsequent large language model.
[0048] In some embodiments, the large language model can parse user-input task instructions and generate corresponding task execution step inference results. In this application, after the visual language adapter converts image features into text features in linguistic form, the large language model can perform further inference and generation operations based on the text features in linguistic form converted from image features. The large language model can not only perform basic question-answering and translation tasks, but also cope with complex remote sensing tasks through multi-turn dialogue and contextual reasoning.
[0049] In some embodiments, Figure 3 This application provides an embodiment of a flowchart for obtaining the reasoning result of the task execution steps corresponding to a first text task instruction. Figure 1 ,like Figure 3 As shown, steps 201 to 205 are used to obtain the task execution step inference results corresponding to the first text task instruction based on the remote sensing image to be processed and the first text task instruction, through a multimodal large model: Step 201: Based on the remote sensing image to be processed, obtain the remote sensing image with adjusted resolution through the data dynamic filling module.
[0050] Step 202: Based on the resolution-adjusted remote sensing image, obtain the segmented image blocks through the embedding module.
[0051] Step 203: Based on the segmented image blocks, obtain the corresponding visual features through a visual encoder.
[0052] Step 204: Based on visual features, obtain the corresponding vector features in the word vector space through a visual language adapter.
[0053] Step 205: Based on vector features and the first text task instruction, obtain the task execution step inference result corresponding to the first text task instruction through the large language model.
[0054] In some embodiments, Figure 4 This application provides a schematic diagram of a process for co-optimization of cognition and perception in a multimodal large model. Figure 1 ,like Figure 4As shown, the data dynamic filling module in step 201 has a preset visual preprocessing strategy. This strategy selects the optimal resolution from multiple candidate resolutions based on the current resolution of the input remote sensing image as the resolution adjustment target for the remote sensing image, and then adjusts the resolution of the remote sensing image based on the resolution adjustment target. The embedding module in step 202 divides the resolution-adjusted remote sensing image into image blocks based on the predefined size of each image block. The visual encoder in step 203 performs an embedding operation on the received image blocks to flatten the image blocks into one-dimensional vectors, and then maps the one-dimensional vectors through linear projection. The vectors are embedded into a fixed-dimensional vector space to obtain embedded vectors. These embedded vectors are then processed by an encoder based on the Transformer framework to obtain corresponding serialized visual features. In step 204, after receiving the serialized visual features, the multilayer perceptron in the visual language adapter maps the serialized visual features to the word vector space through a cross-attention mechanism to obtain the corresponding vector features. In step 205, the large language model performs semantic understanding and planning based on the received vector features and the first text task instructions, and decomposes complex sentences to obtain the corresponding task execution step inference results.
[0055] As can be understood, in this embodiment, firstly, the resolution of the input remote sensing image to be processed is adjusted by a data dynamic filling module to ensure the effectiveness of subsequent image patch segmentation and feature extraction; then, the resolution-adjusted remote sensing image is segmented by an embedding module to facilitate parallel processing and improve overall processing efficiency; subsequently, features are extracted from the segmented image patches by a visual encoder to generate visual features that can be used for cross-modal modeling; then, the visual features are mapped to a word vector space by a visual language adapter, enabling the large language model to understand the image content; finally, the large language model combines image features with task instructions to generate specific execution step inference results. The entire process forms a complete closed loop from image preprocessing to task inference, ensuring the efficient operation and intelligent decision-making capabilities of the multimodal large model.
[0056] In some embodiments, the remote sensing image to be processed in this application includes remote sensing images of multiple modalities, such as optical images acquired by an optical sensor, radar images acquired by a radar sensor, and infrared images acquired by an infrared sensor; the embedding module in this application, matching the modality of the remote sensing image to be processed, includes an optical embedding module, a radar embedding module, and an infrared embedding module. Based on this, embodiments of this application also provide a schematic diagram of the cognitive and perceptual collaborative optimization process of a multimodal large model. Figure 2 ,like Figure 5As shown, the remote sensing images of the corresponding modalities, after resolution adjustment by the data dynamic filling module, are input into the corresponding embedding modules. Furthermore, the three embedding modules for receiving remote sensing images of different modalities share a single visual encoder; that is, the optical embedding module, radar embedding module, and infrared embedding module all input the corresponding divided image blocks into the same visual encoder for visual feature extraction.
[0057] It should be understood that, in the embodiments of this application, the segmented image patches can be obtained based on the resolution-adjusted remote sensing image through the embedding module via the following steps 301 to 303: Step 301: Based on the resolution-adjusted optical image, obtain the divided optical image blocks through the optical embedding module.
[0058] Step 302: Based on the radar image with adjusted resolution, obtain the segmented radar image blocks through the radar embedding module.
[0059] Step 303: Based on the resolution-adjusted infrared image, obtain the divided infrared image blocks through the infrared embedding module.
[0060] It is understood that, in the embodiments of this application, by inputting the obtained optical images, radar images and infrared images into the corresponding embedding modules respectively, and performing feature extraction and encoding in the same visual encoder, multimodal information can be fused in a unified vector space, thereby improving the overall accuracy of recognition and analysis, enhancing the robustness and generalization ability of the multimodal large model in complex environments, and thus supporting a wider range of application scenarios.
[0061] In some embodiments, the resolution-adjusted remote sensing image can be obtained based on the remote sensing image to be processed through a data dynamic filling module via steps 401 to 403 as follows: Step 401: The data dynamic filling module determines the current resolution of the remote sensing image to be processed and adjusts it to multiple filling areas corresponding to multiple preset candidate resolutions.
[0062] For example, multiple preset candidate resolutions include 128*128 resolution, 512*512 resolution and 1024*1024 resolution. The current resolution of the remote sensing image to be processed is 100*100. The first filling area, the second filling area and the third filling area required to adjust the current resolution of 100*100 to 128*128 resolution, 512*512 resolution and 1024*1024 resolution are calculated respectively.
[0063] Step 402: The data dynamic filling module selects the candidate resolution corresponding to the smallest filling area among multiple filling areas as the target resolution.
[0064] For example, it is obvious that the first fill area is the smallest, so the resolution (128*128) corresponding to the first fill area is selected as the target resolution.
[0065] Step 403: The data dynamic filling module adjusts the resolution of the remote sensing image to be processed based on the target resolution to obtain the remote sensing image after resolution adjustment.
[0066] For example, the current resolution (100*100) of the remote sensing image to be processed is adjusted to the target resolution (128*128) to obtain a remote sensing image with a resolution of 128*128.
[0067] It is understood that in the embodiments of this application, the data dynamic filling module dynamically selects an optimal target resolution from a number of pre-set candidate resolutions based on the current resolution of the input remote sensing image, and then adjusts the resolution of the remote sensing image based on the selected target resolution. This adjustment method can reduce redundant information introduced during image processing, thereby improving the accuracy and efficiency of image processing, and further enhancing the perception and cognitive collaboration capabilities of multimodal large models in remote sensing tasks.
[0068] In some embodiments, the training dataset further includes an image-text pairing dataset; the image-text pairing dataset includes a first remote sensing image and descriptive text corresponding to the first remote sensing image; wherein, the image-text pairing dataset is used to train the visual language adapter to align images and text.
[0069] It is understood that, in this embodiment, the descriptive text is natural language written to correspond to the content contained in the first remote sensing image. This aims to help the visual language adapter in the multimodal large model understand the semantic content of the image, thereby enabling the visual language adapter to fully learn the mapping rules between the remote sensing image and the text language. During the alignment training process, other modules are frozen to ensure their parameters remain unchanged. Then, through repeated training and gradual optimization of the visual language adapter's parameter settings, the visual language adapter can more accurately capture key information in the image and transform this key information into understandable language expressions, thus achieving efficient cross-modal understanding and reasoning.
[0070] In the embodiments of this application, since the multimodal large model can understand complex text semantics in practical applications, even when complex text task instructions are input into it, it can leverage its powerful semantic decomposition and planning capabilities to break down the complex text task instructions into corresponding task execution reasoning steps, thereby achieving multi-step remote sensing task processing. The reason why the multimodal large model of this application possesses semantic decomposition and planning capabilities is that it incorporates the execution steps of text task instructions into the training dataset, enabling the multimodal large model to accurately decompose complex sentences. Furthermore, the multimodal large model of this application also dynamically adjusts the resolution and divides the input remote sensing images into image patches, accelerating the visual encoder's processing of remote sensing images while ensuring the representation accuracy of the remote sensing images.
[0071] Another embodiment of this application provides a training method for a multimodal large model, which is applied to a training device or electronic device for a multimodal large model.
[0072] Figure 6 A schematic diagram of the implementation process of the multimodal large model training method provided in the embodiments of this application. Figure 1 .like Figure 6 As shown, the method includes the following steps 501 to 502: Step 501: Based on the training dataset, obtain the task execution step prediction results through the initial multimodal large model; wherein, the training dataset includes a remote sensing image dataset and a task instruction dataset; the remote sensing image dataset includes multimodal remote sensing data; the task instruction dataset includes a second text task instruction and the task execution steps corresponding to the second text task instruction.
[0073] Step 502: Based on the prediction results of the task execution steps and the task execution steps, correct the initial multimodal large model to obtain the multimodal large model; wherein, the multimodal large model is used to obtain the task execution step inference results corresponding to the first text task instruction through the remote sensing image to be processed and the first text task instruction.
[0074] It is understood that, in this embodiment, by constructing a remote sensing image dataset containing multimodal remote sensing data, a second text task instruction, and a task instruction dataset containing the corresponding task execution steps of the second text task instruction, the initial multimodal large model's ability to understand complex tasks is improved. Then, by comparing the predicted task execution steps with the actual task execution steps input during training, the initial multimodal large model is corrected, enabling it to continuously learn how to more accurately parse task instructions and generate execution steps that are closer to reality. The model formed after correction is called the multimodal large model, which possesses stronger task generalization ability and higher execution accuracy. Furthermore, based on the remote sensing image to be processed and the first text instruction, the corrected multimodal large model can obtain more realistic task execution step inference results.
[0075] It should be understood that, in the embodiments of this application, the multimodal remote sensing data includes remote sensing images of multiple modes, such as optical images obtained by optical sensors, radar images obtained by radar sensors, and infrared images obtained by infrared sensors; the first text task instruction and the second text task instruction are instructions issued based on the remote sensing task.
[0076] In some embodiments, the initial multimodal large model includes a data padding module for adjusting the resolution of remote sensing images, an embedding module for dividing remote sensing images into image patches, a visual encoder for extracting visual features of remote sensing images, a visual language adapter for mapping visual features to a word vector space, and a large language model for processing text tasks.
[0077] In this embodiment, the resolution of the input remote sensing image is dynamically adjusted by the data filling module, ensuring the quality of the remote sensing image during training while reducing unnecessary redundant information, thereby improving the processing efficiency of subsequent remote sensing images. The partitioning of the embedding module ensures that the visual encoder better captures the multi-scale features of the remote sensing image during training. Through visual feature extraction by the visual encoder, the spatial hierarchical relationships of the remote sensing image can be automatically learned during training, generating visual feature vectors with strong semantic expression capabilities, significantly improving the initial multimodal large model's ability to handle complex remote sensing tasks. The mapping of visual features to word vector space by the visual language adapter enables the large language model to fully understand and utilize the visual features of the remote sensing image during training. The large language model's parsing of text task instructions and task execution steps during training allows it to understand and parse complex text task instructions, thereby outputting execution steps that meet the task requirements.
[0078] In some embodiments, the training dataset further includes an image-text pairing dataset; the image-text pairing dataset includes a first remote sensing image and descriptive text corresponding to the first remote sensing image; and the visual language adapter in the initial multimodal large model is aligned and trained through the following steps 601 to 602: Step 601: Based on the first remote sensing image, obtain the descriptive predicted text through a visual language adapter.
[0079] Step 602: Correct the visual language adapter based on the predicted descriptive text and the descriptive text.
[0080] It is understood that, in this embodiment of the application, by introducing an image-text pairing dataset into the training dataset and optimizing the visual language adapter based on the image-text pairing dataset, the ability of the visual language adapter to capture complex semantic relationships corresponding to remote sensing images can be significantly improved during the training phase, thereby achieving accurate alignment between images and text, and thus enabling more accurate completion of remote sensing tasks.
[0081] In the embodiments of this application, since the training dataset used during the initial training of the multimodal large model includes not only remote sensing images and text task instructions, but also task execution steps corresponding to the text task instructions, the initial multimodal large model possesses semantic decomposition and planning capabilities during training. Furthermore, during training, the initial multimodal large model is corrected based on the predicted task execution steps and the actual task execution steps, enabling it to output realistic task execution steps during training. This avoids the need to design specific training datasets for different task scenarios when applying it to different tasks. Moreover, an image-text pairing dataset is added to the training dataset, and the initial multimodal large model is further corrected based on the predicted descriptive text, allowing it to continuously optimize its semantic understanding of image content, thereby enhancing its overall inference performance.
[0082] Based on the above embodiments, the following examples describe possible implementation schemes of the model training method for remote sensing applications of one or more of the above embodiments.
[0083] In key areas such as regional integrated situational awareness and emergency response to major disasters, including high-resolution optical imaging and synthetic aperture radar (SAR), Radar (SAR) data and multimodal remote sensing data from thermal infrared remote sensing are increasingly becoming core data support for complex scene cognition and intelligent decision-making in these key fields. Traditional remote sensing intelligent analysis techniques are mainly based on discriminative models specifically designed for single tasks, such as target detection frameworks or semantic segmentation algorithms based on deep convolutional neural networks. While these models can achieve good performance in specific scenarios and preset tasks, their inherent limitations become increasingly apparent as application scenarios become more complex and task requirements diversify. First, these models typically adopt a "one task, one model" design philosophy, with the model architecture and parameter space strictly limited to the specific task scope, making it unable to flexibly adapt to dynamically changing task requirements in actual applications. Second, when it is necessary to expand to new task types or target categories, it is often necessary to redesign the model architecture or perform full fine-tuning of parameters, which not only leads to huge consumption of computing resources but also causes the model deployment and maintenance costs to increase exponentially. In addition, existing methods are limited to two-dimensional planar analysis in terms of spatial understanding, and cannot fully utilize the three-dimensional spatial relationships contained in remote sensing data.
[0084] In recent years, with breakthroughs in multimodal large language model technology, intelligent analysis methods based on generative pre-training frameworks have brought revolutionary opportunities to the field of remote sensing. These models, by constructing a unified visual-language joint representation space, have initially achieved cross-modal semantic alignment and task generalization capabilities. However, current multimodal large models applied in remote sensing still face many challenges in practical applications. In terms of cognitive capabilities, existing models mostly adopt an end-to-end single-round reasoning paradigm, lacking the ability to decompose and plan complex tasks, making them ill-suited for comprehensive tasks requiring multi-step collaboration. Simultaneously, in terms of spatial understanding, existing remote sensing multimodal large language models cannot establish three-dimensional spatial relationships between targets.
[0085] In the research and analysis of related technologies, this application has found the following defects in the related technologies: 1. Traditional remote sensing intelligent analysis models are usually designed based on a single task and lack the ability to flexibly respond to dynamic task requirements.
[0086] 2. When expanding to new tasks, the existing model needs to be redesigned or fully fine-tuned, which leads to increased consumption of computing resources and maintenance costs.
[0087] 3. Current multimodal large models lack the ability to decompose and plan complex tasks, making it difficult to handle comprehensive tasks that require multi-step collaboration.
[0088] 4. Existing methods are limited to two-dimensional planar analysis in terms of spatial understanding and cannot fully utilize the three-dimensional spatial relationships of remote sensing data.
[0089] To address the aforementioned issues, this application provides a remote sensing cognition and perception collaborative optimization method based on a multimodal large model. This method can understand the input characteristics of multi-source remote sensing data, such as optical, synthetic aperture radar, and infrared data, and, combined with the different needs of the perception and cognition links in the task chain, achieve dynamic intelligent output and collaborative decision optimization in multi-task scenarios. The remote sensing cognition and perception collaborative optimization method based on a multimodal large model can play an important role in national security defense, maritime traffic management, civil aviation flight safety, maritime and air search and rescue, environmental monitoring, and combating illegal activities. In this embodiment, the research object is remote sensing images taken by extraterrestrial satellites containing key targets such as aircraft and ships. In practical applications, the model needs to understand the key semantic information in multi-source remote sensing images (optical, SAR, infrared, etc.) according to task instructions in natural language form, and output the results of corresponding perception and / or cognition tasks based on task requirements.
[0090] Figure 7 A schematic diagram of the implementation process of a remote sensing cognition and perception collaborative optimization method based on a multimodal large model provided in this application embodiment. Figure 2 ,like Figure 7 As shown, the method includes the following steps 701 to 703: Step 701: Construct the dataset and perform corresponding annotations (data preparation and annotation).
[0091] A training dataset (multi-task remote sensing dataset) for remote sensing tasks was constructed and standardized in a unified format to prepare and annotate the data. The focus was on building a multi-task dataset covering typical remote sensing tasks, primarily categorized into perception and cognitive tasks. Perception tasks include basic visual processing tasks such as target detection, target classification, spatial relationship understanding, and image description; cognitive tasks include higher-level reasoning tasks such as instruction decomposition, behavioral decision-making, and task planning. During data construction, existing open-source remote sensing data resources were collected and expanded, their original structures and annotation methods were analyzed, and a unified format conversion was performed on different data sources. Finally, all remote sensing images and task data were uniformly converted into a standardized question-and-answer pair format.
[0092] Step 702: Construct the initial multimodal large model (construction of a multimodal large model for remote sensing cognition and perception collaborative optimization).
[0093] The initial multimodal large model constructed mainly includes: a data dynamic filling module, an embedding module, a multimodal high-resolution remote sensing image visual encoder, a visual language adapter (visual-language semantic processing module), and a large language model, in order to realize the construction of a multimodal large model oriented towards the collaborative optimization of remote sensing cognition and perception.
[0094] The visual encoder is used to extract and encode features from input multi-source remote sensing image data, supporting image inputs from different modalities such as optical, SAR, and infrared. The visual encoder is obtained through large-scale image pre-training and possesses powerful feature extraction and processing capabilities. It can employ, but is not limited to, visual models based on the Transformer architecture such as Sigmoid Loss for Language Image Pre-Training (SigLIP) and SAM-B, combined with dynamic block partitioning strategies and modality-adaptive embedding structures, to achieve consistent encoding and efficient compressed representation of images at different scales and modalities.
[0095] The visual-language semantic processing module transforms visual features into the dimensional space of a large language model while preserving the semantic invariance of the visual features. Structurally, it includes, but is not limited to, single-layer linear mapping layers, multilayer perceptrons, and Q-Former structures.
[0096] Large language models are pre-trained on massive amounts of text data, possessing the ability to understand, generate, and reason about natural language, and can map complex linguistic information into useful numerical representations. Large language models can perform basic language tasks (such as text classification, question answering, and translation), as well as handle the demands of complex scenarios through multi-turn dialogue and contextual reasoning. Large language models include, but are not limited to, Large Language Model MetaAI (LLaMA), with corresponding sizes of 7B, 13B, and 65B; Vicuna, with corresponding sizes of 7B and 13B; Open Pretrained Transformer (OPT), with corresponding sizes of 125M, 350M, 1.3B, 6.7B, 13B, and 175B; and DeepSeekMoE, with corresponding sizes of 3B, 16B, and 27B, among other open-source large language models.
[0097] Step 703: Train the initial multimodal large model based on the constructed training dataset (multi-stage training of the basic model).
[0098] Based on the training dataset constructed for multiple scenarios in step 701, the initial multimodal large model constructed in step 702 is trained in multiple stages to obtain the multimodal large model. The multiple stages include a pre-training stage and a supervised fine-tuning stage.
[0099] Step 704: Input the remote sensing image to be processed (unlabeled high-resolution remote sensing image), and combine it with the first text task instruction (task instruction) to enable the multimodal large model to complete multi-type inference tasks.
[0100] Unlabeled high-resolution remote sensing images and task instructions provided in text form are input into the obtained multimodal large-scale model, guiding the model to perform integrated perception and cognition task reasoning. Task instructions may include, but are not limited to, target category recognition, target location localization, and path planning strategies. Based on understanding the unlabeled high-resolution remote sensing images and task instructions (text) input, the multimodal large-scale model automatically completes cross-modal task parsing and response, outputting task execution step reasoning results (structured or natural language answers) that satisfy the task execution (task description).
[0101] Furthermore, the training dataset constructed in step 701 includes: Remote sensing image target detection dataset: This dataset can be an existing publicly available dataset, such as Dota aerial image object recognition (DIOR) or Dataset for Object Detection in Aerial Images (DOTA). Alternatively, it can be a dataset annotated using target detection annotation software, containing remote sensing images and the coordinates and category information of the bounding boxes of key targets in the corresponding images.
[0102] Target description text dataset: This dataset can be based on remote sensing image classification datasets such as the Northwestern Polytechnical University Remote Sensing Image Scene Classification Benchmark (NWPU-RESISC45) and the Aerial Image Dataset (AID). It can also extract labeled target image regions from remote sensing image target detection data and organize them according to categories to construct a dataset for target type recognition, supporting the training of the model's discrimination ability in multi-class target scenes.
[0103] Remote Sensing Image Spatial Relationship Understanding Dataset: This dataset can be constructed by labeling the relative spatial relationships between targets in remote sensing images (such as "ships are near the port", "aircraft are in the center of the runway") based on target detection images in remote sensing images, or by extracting structured labels from artificially generated spatial relationship description texts. It can be used to train models to understand spatial semantic relationships such as relative orientation, distance, and containment between targets.
[0104] Remote Sensing Image Description Dataset: This dataset includes remote sensing images and their corresponding natural language descriptions. The descriptions cover the main targets, scene types, and their combined features in the remote sensing images. Data can be obtained from existing remote sensing image description datasets such as the Remote Sensing Image Captioning Dataset (RSICD) and UCM-Captions. Alternatively, descriptions can be manually written or automatically generated using weak supervision and then manually revised to improve accuracy and diversity.
[0105] Instruction Decomposition Dataset: This dataset contains second-level text task instructions (natural language instructions) and their corresponding decomposition steps, aiming to guide models in understanding the structural composition of complex instructions. For example, "describe information within this region" is decomposed into "locate the target region," "extract the target," "analyze the relationships between targets," and "summarize." The data can be obtained by constructing multi-step task instructions and annotating the decomposition process. This dataset is primarily based on natural language instructions in two-dimensional spatial coordinates and their corresponding decomposition steps.
[0106] Behavioral Decision Dataset: This dataset is used to train models to make intelligent decisions based on remote sensing images and task instructions. The data format includes multimodal remote sensing data (input images), second-text task instructions (task text), and corresponding task execution steps (decision results, such as "progress along building A", "turn right in area B", "reach target location C", etc.). The data can be combined with actual task simulation scripts and expert experience to construct typical decision scenarios in remote sensing tasks, and corresponding labeled response strategies can be implemented. This dataset is mainly composed of remote sensing data in three-dimensional spatial coordinates, second-text task instructions in three-dimensional spatial coordinates, and task execution steps in three-dimensional spatial coordinates. The second-text task instructions and task execution steps are mainly task execution and task execution steps across remote sensing images, respectively. For example, the second-text task instruction is the path planning for a ship from the starting point M in remote sensing image A to the destination N in remote sensing image B, and the task execution steps are the specific path planning strategies (e.g., first from the starting point M to the target point J in remote sensing image A, then from the target point J to the target point K in remote sensing image B, and finally from the target point K to the destination N).
[0107] Task Planning Dataset: This dataset contains multi-step task execution processes, including remote sensing data, second-text task instructions (location of key surrounding objects in the target area from the remote sensing data, target location identification, and task initiation conditions), and corresponding task execution steps (operation paths) for the second-text task instructions. It is used to train models to possess planning and scheduling capabilities. Data construction can refer to typical datasets such as CityNAV, and is completed through manual arrangement or annotation using rule-based generation tools. This dataset is mainly composed of remote sensing data in three-dimensional spatial coordinates, second-text task instructions in three-dimensional spatial coordinates, and task execution steps in three-dimensional spatial coordinates. The second-text task instructions and task execution steps primarily represent the task execution of blurred targets in remote sensing images and the task execution steps required to achieve the blurred targets, respectively.
[0108] Furthermore, the data dynamic filling module in the multimodal large model for remote sensing cognition and perception collaborative optimization constructed in step 702 is used to adjust the resolution of the input remote sensing images (including remote sensing images from the training phase and the actual application phase). The adjustment strategy is a visual preprocessing strategy based on dynamic stitching to adapt to different aspect ratios commonly found in remote sensing images. This visual preprocessing strategy includes the following steps: First, multiple candidate resolutions are preset; then, the current resolution of the input remote sensing image is adjusted to the multiple fill areas corresponding to the multiple candidate resolutions; second, the candidate resolution corresponding to the smallest fill area among the multiple fill areas is selected as the target resolution; finally, the resolution of the input remote sensing image is adjusted based on the determined target resolution.
[0109] For example, the visual preprocessing strategy is as follows: Preset a set of candidate resolutions:
[0110] Where m:n represents the aspect ratio of the input remote sensing image (candidate image); For any input remote sensing image (height H, width W), we iterate through C R Calculate the minimum fill area required to adjust the current resolution of the remote sensing image to the traversed resolutions among all candidate resolutions. Select the resolution with the smallest fill area As the target resolution (final adjustment target); when the selected target resolution is ( )hour, and It is 3; The resolution of the input remote sensing image is adjusted based on the selected target resolution.
[0111] Furthermore, the embedded module in the multimodal large model for remote sensing cognition and perception co-optimization constructed in step 702 is used to divide the resolution-adjusted remote sensing image into image patches; for example, each remote sensing image is divided into... A local image patch (each patch is 1000 pixels in size) ) and a global thumbnail tile.
[0112] The visual encoder in the multimodal large model for remote sensing cognition and perception co-optimization constructed in step 702 is used to analyze this... Encode each image block to generate a block with a size of [size missing]. (For example, The visual embeddings are M in dimension (e.g., 1152) of each embedding. To adapt to the representational characteristics of different modalities in remote sensing images, the differentiated visual embedding layers are readjusted after freezing the visual encoder to process optical, SAR and infrared images respectively.
[0113] After visual embedding extraction, a pixel-level (e.g.,) process is further performed once within each image block. A pixel-level shuffling operation that changes the visual markers of each piece from their original size (e.g., ...). Compress to the appropriate size (e.g., ), thereby obtaining the corresponding number of tags (e.g., total). 196 tags).
[0114] After the visual language adapter in the multimodal large model for remote sensing cognition and perception co-optimization constructed in step 702 receives the complete visual input sequence containing 196 labels, the sequence will be projected into the word vector space (embedding space) of the large language model through two layers of multilayer perceptrons in the visual language adapter, thereby achieving efficient connection between visual and language representation.
[0115] The visual feature sequence mapped by the visual language adapter and the text task instructions are used as input to the large language model to achieve the fusion understanding and joint reasoning of multimodal information.
[0116] Furthermore, the multimodal large-scale model for remote sensing cognition and perception co-optimization constructed in step 702 is trained using the training dataset constructed in step 701. After training with the training dataset constructed in step 701, a large-scale image-text pairing dataset is generated using the large language model, and this image-text pairing dataset is used as part of the training dataset to perform alignment training (alignment operation) on the visual language adapter. The core objective of the alignment training stage is to train the visual language adapter module in the middle of the model to effectively narrow the gap between image features and text semantics, so that the adapter parameters can learn how to more accurately integrate text information to guide the understanding of image content. The dataset used in the pre-training stage integrates image-text alignment data and plain text corpora, aiming to maintain the model's image-text understanding ability while taking into account its performance in plain text tasks, achieving a balanced improvement in multi-task capabilities.
[0117] Furthermore, while maintaining the parameter structure of the multimodal large model for remote sensing cognition and perception co-optimization after training using the training dataset, a supervised fine-tuning dataset is constructed by combining the training dataset built in step 701 with high-quality internal question-answering pairs. The supervised fine-tuning stage of training is then carried out on the multimodal large model for remote sensing cognition and perception co-optimization. The supervised fine-tuning dataset includes a second remote sensing image (selected from the training dataset in step 701), a question set for the second remote sensing image, and an answer set for the question.
[0118] During training, a structured annotation design is adopted for text tasks involving spatial positioning information. This involves adding 2D and 3D spatial coordinates as labels to the training dataset corresponding to the task type. For tasks based on 2D planar coordinates, such as object detection and spatial relationship analysis, and for tasks involving 3D spatial coordinates, such as path planning and scheduling strategies, specific guiding words are used to elicit corresponding labels. All coordinate information is uniformly normalized to a numerical range of 0 to 999 based on image resolution. This structured label (coordinate information) is embedded in the overall output in text form. During training, the coordinate values are optimized by calculating the text cross-entropy loss function, thus maintaining consistency with the calculation method of other natural language answers and achieving unified modeling and parallel training.
[0119] Furthermore, during model training, based on the task execution step prediction results (output results) generated by the large language model, the cross-entropy loss function is used to calculate the difference between the task execution step prediction results and the task execution steps (true labels) in order to optimize model parameters and improve the accuracy and robustness of the generated results.
[0120] It is understood that the remote sensing cognition and perception collaborative optimization method based on a multimodal large model provided in this application can flexibly process multimodal remote sensing images such as optical, SAR, and infrared images according to various input natural language commands. It possesses a deep understanding of target attributes, spatial relationships, and environmental semantics within the images, effectively supporting diverse remote sensing perception and cognition tasks. This meets the needs of a wide range of applications, including national defense, environmental monitoring, and traffic management.
[0121] In the embodiments of this application, compared with traditional remote sensing basic models that usually only have one capability of perception or cognition, the remote sensing cognition and perception collaborative optimization method of multimodal large model proposed in this application can construct a sensing and control integrated basic model system, flexibly handle different types of perception and control tasks, and thus be more suitable for stable operation in complex display environments.
[0122] In the embodiments of this application, compared to traditional perception-based remote sensing basic models that can typically only handle a single modality, the multimodal large-scale remote sensing cognition and perception collaborative optimization method proposed in this application can flexibly handle optical, SAR, and infrared data, effectively perform domain-specific feature extraction for different modalities, share the visual encoder, and fully extract features of different modalities by only fine-tuning a small portion of the visual encoding.
[0123] In the embodiments of this application, compared to traditional perception-based remote sensing basic models which typically have large parameters and high computational resource consumption, making them difficult to deploy efficiently in resource-constrained environments, the multimodal large-scale remote sensing cognition and perception collaborative optimization method proposed in this application optimizes computational efficiency while reducing computational costs, ensuring efficient operation under various complex tasks and making it suitable for resource-limited application scenarios.
[0124] In the embodiments of this application, compared to the limitations of traditional perception-based remote sensing basic models that require retraining or fine-tuning when dealing with new tasks, the multimodal large-scale remote sensing cognition and perception collaborative optimization method proposed in this application can seamlessly switch between multiple tasks, flexibly respond to the needs of different target categories or task types, and significantly improve the efficiency and adaptability of remote sensing image analysis.
[0125] The remote sensing cognition and perception collaborative optimization method based on a multimodal large model provided in this application adds task execution steps corresponding to text task instructions to the training dataset. This enables the multimodal large model trained on the training dataset to output corresponding task execution inference steps based on the input remote sensing image and text task instructions. Thus, the multimodal large model provided in this application has the ability to understand and plan complex text, thereby enabling the multimodal large model to be effectively applied to complex text task scenarios.
[0126] The multimodal large model training method provided in this application adds task execution steps corresponding to text task instructions to the training dataset, enabling the initial multimodal large model to have text segmentation and planning capabilities during the training process. Furthermore, the parameters of the initial multimodal large model are continuously optimized based on the predicted results of the inferred task execution steps and the actual task execution steps, thereby continuously improving the semantic understanding capability of the initial multimodal large model. This ensures that the initial multimodal large model can adapt to training datasets of various task scenarios, avoiding the need to design separate training datasets for specific scenarios.
[0127] The initial multimodal large-scale model provided in this application can understand the input characteristics of multi-source remote sensing data such as optical, SAR, and infrared data. Combining the different needs of the perception and cognition stages in the task chain, it achieves dynamic intelligent output and collaborative decision optimization in multi-task scenarios. A large-scale remote sensing image dataset is constructed and used for training. This dataset contains high-resolution images and their corresponding two-dimensional and three-dimensional spatial coordinates, spatial location coordinates between targets, etc., enhancing the model's understanding and reasoning ability regarding remote sensing images. A visual preprocessing strategy based on dynamic stitching is introduced, which can dynamically adjust the receptive field and optimize the image patch segmentation method according to the actual input image size, effectively improving the processing efficiency and performance of high-resolution remote sensing images. To improve feature accuracy, a unified encoding mechanism (sharing a single visual encoder) for heterogeneous modal features from multiple sources of remote sensing is provided, focusing on the feature mapping relationships between typical modalities such as optical, SAR, and infrared, and constructing a shared high-dimensional feature space. For decision-making and scheduling tasks in remote sensing scenarios, task expansion and custom data creation are further carried out in conjunction with existing relevant datasets to construct complex task scenarios covering multiple objectives and constraints, thereby improving the model's reasoning and decision-making capabilities in high-complexity tasks. A multi-stage pre-training strategy is adopted, combined with feature alignment and instruction tuning methods, to solve the modal gap and domain shift problems in the migration of large models from natural scenes to the remote sensing domain, and improve the model's transferability and generalization performance in remote sensing multi-task chains.
[0128] It should be noted that although the steps of the method in this application are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps; or steps from different embodiments may be combined into a new technical solution.
[0129] Based on the foregoing embodiments, this application provides a remote sensing cognition and perception collaborative optimization device based on a multimodal large model. The device includes various modules and units included in each module, which can be implemented by a processor; of course, it can also be implemented by specific logic circuits. In the implementation process, the processor can be an AI acceleration engine (such as NPU), GPU, central processing unit (CPU), microprocessor (MPU), digital signal processor (DSP), or field programmable gate array (FPGA), etc.
[0130] Figure 8 A schematic diagram of a remote sensing cognition and perception collaborative optimization device based on a multimodal large model provided in this application embodiment is shown below. Figure 8 As shown, the remote sensing cognition and perception collaborative optimization device 800 based on a multimodal large model includes a processing unit 801 and an execution unit 802, wherein: The processing unit 801 is configured to obtain the task execution step inference result corresponding to the first text task instruction based on the remote sensing image to be processed and the first text task instruction through a multimodal large model; wherein, the multimodal large model is obtained by training a training dataset; the training dataset includes a remote sensing image dataset and a task instruction dataset; the remote sensing image dataset includes multimodal remote sensing data; the task instruction dataset includes a second text task instruction and the task execution steps corresponding to the second text task instruction.
[0131] Execution unit 802 is configured to perform remote sensing task processing based on the reasoning results of the task execution steps.
[0132] Based on the foregoing embodiments, this application also provides a training device for a multimodal large model, the structural schematic diagram of which is shown below. Figure 9 As shown, the training device 900 for a multimodal large model includes a training unit 901 and a correction unit 902, wherein: Training unit 901 is configured to obtain task execution step prediction results based on the training dataset and an initial multimodal large model; wherein, the training dataset includes a remote sensing image dataset and a task instruction dataset; the remote sensing image dataset includes multimodal remote sensing data; the task instruction dataset includes a second text task instruction and the task execution steps corresponding to the second text task instruction; The correction unit 902 is configured to correct the initial multimodal large model based on the prediction results of the task execution steps and the task execution steps to obtain a multimodal large model; wherein, the multimodal large model is used to obtain the task execution step inference results corresponding to the first text task instruction through the remote sensing image to be processed and the first text task instruction.
[0133] The descriptions of the above device embodiments are similar to those of the above method embodiments, and have similar beneficial effects. For technical details not disclosed in the device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0134] It should be noted that the module division in the embodiments of this application is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods. Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, exist as separate physical units, or have two or more units integrated into one unit. The integrated units can be implemented in hardware, as software functional units, or a combination of software and hardware.
[0135] It should be noted that, in the embodiments of this application, if the above methods are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device to execute all or part of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware and software combination.
[0136] This application provides an electronic device. Figure 10 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application, such as... Figure 10 As shown, the electronic device 2000 includes a memory 2001 and a processor 2002. The memory 2001 stores a computer program that can run on the processor 2002. When the processor 2002 executes the program, it implements the steps in the method provided in the above embodiments.
[0137] It should be noted that the memory 2001 is configured to store instructions and applications executable by the processor 2002, and can also cache data to be processed or already processed (e.g., image data, audio data, voice communication data and video communication data) in the processor 2002 and various modules in the electronic device 2000. It can be implemented by flash memory or random access memory (RAM).
[0138] This application also provides a computer-readable storage medium for storing computer programs.
[0139] Optionally, the computer-readable storage medium can be applied to the electronic device in the embodiments of this application, and the computer program causes the processor to execute the various methods of the embodiments of this application, which will not be described in detail here for the sake of brevity.
[0140] This application also provides a computer program product, including computer program instructions.
[0141] Optionally, the computer program product can be applied to the electronic device in the embodiments of this application, and the computer program instructions cause the processor to execute the various methods in the embodiments of this application, which will not be described in detail here for the sake of brevity.
[0142] This application also provides a computer program.
[0143] Optionally, the computer program can be applied to the electronic device in the embodiments of this application. When the computer program runs on the processor, it causes the processor to execute the various methods of the embodiments of this application. For the sake of brevity, these will not be described in detail here.
[0144] It should be noted that the descriptions of the electronic devices, storage media, computer program products, and computer program embodiments above are similar to the descriptions of the method embodiments above, and have similar beneficial effects. For technical details not disclosed in the electronic devices, storage media, computer program products, and computer program embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0145] It should be understood that the phrases "one embodiment," "an embodiment," or "some embodiments" throughout the specification mean that a specific feature, structure, or characteristic related to an embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment," "in one embodiment," or "in some embodiments" appearing throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely descriptive and do not represent the superiority or inferiority of the embodiments. The descriptions of the various embodiments above tend to emphasize the differences between the various embodiments; their similarities or commonalities can be referred to mutually, and for the sake of brevity, they will not be repeated here.
[0146] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three kinds of relationships. For example, object A and / or object B can represent three situations: object A exists alone, object A and object B exist simultaneously, and object B exists alone.
[0147] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0148] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple modules or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or modules can be electrical, mechanical, or other forms.
[0149] The modules described above as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules. They may be located in one place or distributed across multiple network units. Some or all of the modules may be selected to achieve the purpose of this embodiment according to actual needs.
[0150] In addition, each functional module in the various embodiments of this application can be integrated into one processing unit, or each module can be a separate unit, or two or more modules can be integrated into one unit; the integrated modules can be implemented in hardware or in the form of hardware plus software functional units.
[0151] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.
[0152] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device to execute all or part of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROMs, magnetic disks, or optical disks.
[0153] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0154] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0155] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.
[0156] The above are merely embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A remote sensing cognition and perception collaborative optimization method based on a multimodal large model, characterized in that, The method includes: Based on the remote sensing image to be processed and the first text task instruction, a multimodal large model is used to obtain the task execution step inference result corresponding to the first text task instruction; wherein, the multimodal large model is obtained by training a training dataset; the training dataset includes a remote sensing image dataset and a task instruction dataset; the remote sensing image dataset includes multimodal remote sensing data; the task instruction dataset includes a second text task instruction and the task execution steps corresponding to the second text task instruction; Remote sensing task processing is performed based on the reasoning results of the aforementioned task execution steps.
2. The remote sensing cognition and perception collaborative optimization method according to claim 1, characterized in that, The multimodal large model includes: a data dynamic filling module for adjusting the resolution of remote sensing images, an embedding module for dividing remote sensing images into image patches, a visual encoder for extracting visual features of remote sensing images, a visual language adapter for mapping visual features to word vector space, and a large language model for processing text tasks.
3. The remote sensing cognition and perception collaborative optimization method according to claim 2, characterized in that, Based on the remote sensing image to be processed and the first text task instruction, the process uses a multimodal large model to obtain the task execution step inference result corresponding to the first text task instruction, including: Based on the remote sensing image to be processed, the resolution-adjusted remote sensing image is obtained through the data dynamic filling module; Based on the remote sensing image with adjusted resolution, the embedded module is used to obtain segmented image blocks. Based on the segmented image blocks, the corresponding visual features are obtained through the visual encoder; Based on the aforementioned visual features, the corresponding vector features in the word vector space are obtained through the visual language adapter. Based on the vector features and the first text task instruction, the task execution step inference result corresponding to the first text task instruction is obtained through the large language model.
4. The remote sensing cognition and perception collaborative optimization method according to claim 3, characterized in that, The remote sensing images to be processed include optical images acquired by an optical sensor, radar images acquired by a radar, and infrared images acquired by an infrared sensor; the embedding module includes an optical embedding module, a radar embedding module, and an infrared embedding module. The remote sensing image adjusted based on the resolution is used to obtain segmented image patches through the embedding module, including: Based on the resolution-adjusted optical image, the optical embedding module is used to obtain segmented optical image blocks; Based on the radar image with adjusted resolution, the radar embedding module is used to obtain segmented radar image blocks; Based on the resolution-adjusted infrared image, the infrared embedding module is used to obtain the divided infrared image blocks.
5. The remote sensing cognition and perception collaborative optimization method according to claim 3 or 4, characterized in that, The process of obtaining a resolution-adjusted remote sensing image based on the remote sensing image to be processed through the data dynamic filling module includes: The data dynamic filling module determines the current resolution of the remote sensing image to be processed and adjusts it to multiple filling areas corresponding to multiple preset candidate resolutions; The data dynamic filling module takes the candidate resolution corresponding to the smallest filling area among the multiple filling areas as the target resolution; The data dynamic filling module adjusts the resolution of the remote sensing image to be processed based on the target resolution to obtain the remote sensing image with adjusted resolution.
6. The remote sensing cognition and perception collaborative optimization method according to any one of claims 2, characterized in that, The training dataset also includes an image-text pairing dataset; the image-text pairing dataset includes a first remote sensing image and descriptive text corresponding to the first remote sensing image; wherein, the image-text pairing dataset is used to train the visual language adapter to align images and text.
7. A training method for a multimodal large model, characterized in that, The method includes: Based on the training dataset, the prediction results of task execution steps are obtained through an initial multimodal large model; wherein, the training dataset includes a remote sensing image dataset and a task instruction dataset; the remote sensing image dataset includes multimodal remote sensing data; the task instruction dataset includes a second text task instruction and the task execution steps corresponding to the second text task instruction; Based on the predicted results of the task execution steps and the task execution steps, the initial multimodal large model is corrected to obtain a multimodal large model; wherein, the multimodal large model is used to obtain the task execution step inference results corresponding to the first text task instruction through the remote sensing image to be processed and the first text task instruction.
8. The training method according to claim 7, characterized in that, The initial multimodal large model includes a data dynamic filling module for adjusting the resolution of remote sensing images, an embedding module for dividing remote sensing images into image patches, a visual encoder for extracting visual features of remote sensing images, a visual language adapter for mapping visual features to word vector space, and a large language model for processing text tasks.
9. The training method according to claim 8, characterized in that, The training dataset also includes an image-text pairing dataset; the image-text pairing dataset includes a first remote sensing image and descriptive text corresponding to the first remote sensing image; Based on the first remote sensing image, the descriptive predicted text is obtained through the visual language adapter; The visual language adapter is corrected based on the predicted descriptive text and the descriptive text.
10. A remote sensing cognition and perception collaborative optimization device based on a multimodal large model, the device comprising: The processing unit is configured to, based on the remote sensing image to be processed and the first text task instruction, obtain the task execution step inference result corresponding to the first text task instruction through a multimodal large model; wherein, the multimodal large model is obtained by training a training dataset; the training dataset includes a remote sensing image dataset and a task instruction dataset; the remote sensing image dataset includes multimodal remote sensing data; the task instruction dataset includes a second text task instruction and the task execution steps corresponding to the second text task instruction; The execution unit is configured to perform remote sensing task processing based on the reasoning results of the task execution steps.
11. A training device for a multimodal large model, the device comprising: The training unit is configured to obtain task execution step prediction results based on the training dataset and an initial multimodal large model; wherein, the training dataset includes a remote sensing image dataset and a task instruction dataset; the remote sensing image dataset includes multimodal remote sensing data; the task instruction dataset includes a second text task instruction and the task execution steps corresponding to the second text task instruction; The correction unit is configured to correct the initial multimodal large model based on the prediction result of the task execution steps and the task execution steps to obtain a multimodal large model; wherein, the multimodal large model is used to obtain the task execution step inference result corresponding to the first text task instruction through the remote sensing image to be processed and the first text task instruction.
12. An electronic device comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the method of any one of claims 1 to 6 or 7 to 9.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the method of any one of claims 1 to 6 or 7 to 9.
14. A computer program product comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the method of any one of claims 1 to 6 or 7 to 9.