Remote sensing reasoning segmentation method and system based on visual language model
By employing a remote sensing inference segmentation method based on a visual language model, and utilizing a geometric regularization injection mechanism and a RAO module to optimize rotation-aware segmentation performance, the problem of target rotation and scale changes in remote sensing images is solved, achieving higher-precision remote sensing image segmentation.
Patent Information
- Application Number
- CN202511135353.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-06-16
- Filing Date
- 2025-08-14
- Publication Date
- 2025-11-28
AI Technical Summary
Existing remote sensing image segmentation technologies are insufficient in segmentation performance when dealing with target rotation and scale changes in remote sensing images. They also lack implicit semantic reasoning and geometric information modeling capabilities, making it difficult to meet the needs of complex remote sensing scenarios.
A remote sensing reasoning segmentation method based on a visual language model is adopted. A reasoning description containing spatial, shape and functional cues is generated through a geometric regularization injection mechanism. An RSReason model is constructed, and the rotation-aware segmentation performance is optimized by using the RAO module and Wasserstein distance. The model is trained and evaluated by combining multi-dimensional indicators.
It significantly improves the segmentation accuracy of targets in any direction in remote sensing images, enhances the model's ability to perceive rotating targets, and achieves more refined and semantically driven target segmentation.
Smart Images

Figure CN121033071A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of remote sensing image segmentation technology, and in particular to a remote sensing inference segmentation method and system based on a visual language model. Background Technology
[0002] With the continuous development of visual language models and large-scale language models, the deep semantic alignment capability between images and natural language is constantly improving, driving the development of language-guided image understanding and segmentation tasks. Among them, Representational Understanding (REC) and Representational Segmentation (RES) are key research directions, widely used in fields such as multimodal perception, human-computer interaction, assisted navigation, and remote sensing image analysis.
[0003] To address this challenge, remote sensing representation expression understanding (RS-REC) and segmentation (RS-RES) tasks have been gradually proposed. For example, Sun et al. constructed the first remote sensing representation dataset, and subsequent datasets like RRSIS and RRSIS-D further expanded the data scale and increased the difficulty. Simultaneously, model structures have also evolved; for instance, RMSIN proposed a multi-scale interaction network to enhance the model's generalization ability. However, these methods largely rely on explicit category or location descriptions, lacking the ability to model implicit semantics such as spatial relationships, geometric attributes, and functional roles. To improve the spatial perception and reasoning capabilities of models, researchers have attempted to introduce multimodal reasoning methods based on large language models. For example, LISA achieves language-driven segmentation through cascaded prompts, while PixelLM and SegLLM introduce encoder-decoder structures for language and images to enhance semantic alignment. Despite some progress, these methods still have many limitations in remote sensing scenarios, especially when dealing with targets with irregular orientations and significant scale variations, where segmentation performance still needs improvement.
[0004] Against this backdrop, this invention proposes the Remote Sensing Reasoning Segmentation Task (RSReasonSeg), which emphasizes establishing reasoning-based semantic associations between images and language to address issues such as ambiguity in target representation and unknown categories in remote sensing images. To support this task, the RSReason benchmark is proposed, which generates complex representations containing spatial, scale, and functional descriptions through a geometric attribute injection mechanism, guiding the model to achieve more refined and semantically driven target segmentation.
[0005] Despite the progress made in current technology, existing methods still have the following limitations in terms of model capabilities and evaluation systems:
[0006] (1) Early remote sensing indexing expression understanding methods such as RSVG introduced the Transformer and BERT architecture, but they were mainly oriented towards explicit category recognition and had limited support for implicit semantic reasoning and geometric information modeling.
[0007] (2) Traditional RS-RES benchmarks such as RRSIS and RRSIS-D focus on explicit target representation and lack description of target rotation attributes and functional semantics, making it difficult to cover the complex reasoning needs in real remote sensing scenarios.
[0008] (3) Some methods based on large language models, such as LISA, PixelLM, and SegLLM, have certain language understanding capabilities, but when applied in the field of remote sensing, they show insufficient perception of changes in target orientation and scale, resulting in a significant decrease in the segmentation accuracy of rotating targets.
[0009] (4) Existing multimodal models such as NExT-Chat and GeoGround are mostly limited to basic recognition or question answering tasks in remote sensing images, lacking systematic modeling for complex geometric structures and spatial distributions;
[0010] In summary, current language-guided segmentation techniques for remote sensing images still have significant shortcomings in terms of expressive power, model architecture adaptability, and the completeness of evaluation dimensions. There is an urgent need for a new benchmark and methodology system that can support implicit semantic reasoning, geometric structure perception, and functional semantic understanding to promote the further development of remote sensing semantic segmentation technology. Summary of the Invention
[0011] The purpose of this invention is to solve the problems of insufficient support for implicit reasoning ability and difficulty in handling rotating targets in the existing technology of remote sensing image segmentation.
[0012] The technical solution adopted by this invention to solve its technical problem is: to provide a remote sensing reasoning segmentation method based on a visual language model, comprising the following steps:
[0013] Image-text-mask pairs are obtained from remote sensing image datasets. Geometric attributes of the masks are extracted, and inferential descriptions containing spatial, shape, and functional cues are generated through a geometric regularization injection mechanism to construct an evaluation dataset.
[0014] The RSReason model is constructed by inputting the evaluation dataset into the RSReason model and training and evaluating the model based on multi-dimensional metrics. During training, the RAO module is used to calculate the RAO loss function for each segmentation example, and the target mask is represented as a 2D Gaussian distribution. The difference between the predicted and the actual mask distribution is calculated based on the Wasserstein distance. The model parameters are updated using this loss function to optimize the rotation-aware segmentation performance.
[0015] The image to be segmented and the text describing the object to be segmented are input into the trained RSReason model for segmentation, and the segmentation mask is obtained as the segmentation result.
[0016] The RSReason model includes:
[0017] A multimodal language module jointly infers image and text to generate text sequences with embedded segmentation tags;
[0018] A visual encoder extracts multi-scale features from an image;
[0019] The mask decoder fuses segmentation markers and multi-scale image features to generate a segmentation mask.
[0020] Preferably, the generation of a reasoning description containing spatial, shape, and functional cues through the geometric regularization injection mechanism includes:
[0021] The geometric regularization injection mechanism approximates the mask as a 2D Gaussian ellipse, uses eigenvalue decomposition of the covariance matrix to derive the principal axis and direction of the ellipse, and converts the geometric attributes into natural language descriptions, including descriptions of position, size and direction, as well as functional role inferences, to generate inferential expressions.
[0022] Preferably, the geometric regularization injection mechanism approximates the mask as a 2D Gaussian ellipse, as follows:
[0023]
[0024] Where x, y represent the coordinates of each pixel in the mask in the Cartesian coordinate system; μ x ,μ y These represent the center positions of the mask on the x-axis and y-axis, respectively; M(x,y) represents the function that generates the binary mask. These represent the variances of the mask on the x-axis and y-axis, respectively; σ xy Σ represents the covariance of the mask in the x and y directions; Σ represents the covariance matrix of the 2D Gaussian distribution into which the mask is converted; G(m) represents the 2D Gaussian distribution into which the mask is converted, where m represents the target binary mask; μ represents the mean of the 2D Gaussian distribution into which the mask is converted.
[0025] Preferably, the multi-dimensional metrics include gIoU, cIoU, and Acc@0.5, which are used to comprehensively measure the segmentation performance of the model on targets of different scales and orientations. The different scales include small, medium, and large objects, and the orientations include targets with obvious rotation and targets without obvious rotation.
[0026] Preferably, the RAO module converts the mask into a 2D Gaussian distribution by calculating the total mass, centroid, variance, and covariance of the mask to construct a covariance matrix, thereby obtaining the position, size, and orientation parameters of the mask and enhancing the model's ability to perceive rotating targets.
[0027] Preferably, the difference between the predicted mask distribution calculated based on Wasserstein distance and the actual mask distribution is expressed as follows:
[0028]
[0029] in, and represents the Gaussian distribution of the predicted and true masks, respectively; Tr represents the trace operation, which calculates the sum of the diagonal elements of the matrix.
[0030] This invention also provides a remote sensing reasoning segmentation system based on a visual language model, comprising:
[0031] The data acquisition module obtains image-text-mask pairs from the remote sensing image dataset, extracts the geometric attributes of the masks, and generates inferential descriptions containing spatial, shape, and functional cues through a geometric regularization injection mechanism to construct the evaluation dataset.
[0032] The model building module constructs the RSReason model, inputs the evaluation dataset into the RSReason model, and trains and evaluates the model based on multi-dimensional metrics. During training, the RAO module is used to calculate the RAO loss function for each segmentation example, represents the target mask as a 2D Gaussian distribution, calculates the difference between the predicted and the true mask distribution based on the Wasserstein distance, and uses this loss function to update the model parameters and optimize the rotation-aware segmentation performance.
[0033] The image segmentation module takes the image to be segmented and the text describing the object to be segmented as input into the trained RSReason model for segmentation and obtains the segmentation mask as the segmentation result.
[0034] The RSReason model includes:
[0035] A multimodal language module jointly infers image and text to generate text sequences with embedded segmentation tags;
[0036] A visual encoder extracts multi-scale features from an image;
[0037] The mask decoder, which fuses segmentation markers and multi-scale image features to generate a segmentation mask, has the following advantages: it can accurately locate and segment targets based on implicit text descriptions, improving the segmentation performance of targets in any direction in remote sensing images.
[0038] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments, but the present invention is not limited to the embodiments. Attached Figure Description
[0039] Figure 1 This is a diagram illustrating the method steps of an embodiment of the present invention;
[0040] Figure 2 This is a detailed flowchart of an embodiment of the present invention;
[0041] Figure 3 This is a system structure diagram of an embodiment of the present invention. Detailed Implementation
[0042] refer to Figure 1 The diagram shown illustrates the method steps of an embodiment of the present invention, including:
[0043] S101: Obtain image-text-mask pairs from remote sensing image datasets, extract the geometric attributes of the masks, generate inferential descriptions containing spatial, shape, and functional cues through a geometric regularization injection mechanism, and construct an evaluation dataset.
[0044] S102, construct the RSReason model, input the evaluation dataset into the RSReason model, and train and evaluate the model based on multi-dimensional indicators; during training, use the RAO module to calculate the RAO loss function for each segmentation example, represent the target mask as a 2D Gaussian distribution, calculate the difference between the predicted and the real mask distribution based on the Wasserstein distance, use this loss function to update the model parameters, and optimize the rotation-aware segmentation performance.
[0045] S103, input the image to be segmented and the text describing the object to be segmented into the trained RSReason model for segmentation, and obtain the segmentation mask as the segmentation result;
[0046] Specifically, participate Figure 2 As shown, the RSReason model includes:
[0047] A multimodal language module jointly infers image and text to generate text sequences with embedded segmentation tags;
[0048] A visual encoder extracts multi-scale features from an image;
[0049] The mask decoder fuses segmentation markers and multi-scale image features to generate a segmentation mask.
[0050] Specifically, the multimodal language module receives the image to be segmented and text describing the segmentation object, and outputs a text sequence embedding segmentation tags. A segmentation tag refers to a special tag in this text sequence, namely... Figure 2 The "<mask>" in the text sequence is extracted separately and input into the mask decoder to generate the mask.
[0051] Specifically, in step S101, image-text-mask pairs are obtained from the remote sensing image dataset. The geometric attributes of the masks are extracted, and an inferential description containing spatial, shape, and functional cues is generated through a geometric regularization injection mechanism to construct an evaluation dataset. Specifically, image-text-mask triples are first collected from the remote sensing image dataset, and then the geometric attributes of the masks, including location, size, shape, and orientation, are extracted. Through a geometric regularization injection mechanism, these geometric attributes are transformed into natural language descriptions, generating inferential descriptions containing spatial, shape, and functional cues. These descriptions are used to convert traditional denotative expressions into inferential expressions, thereby constructing the RSReason dataset for evaluating the model's inference capabilities.
[0052] Specifically, in S102, the RSReason model is used for segmentation. This model includes a Rotation Aware Optimization (RAO) module, which represents the target mask as a 2D Gaussian distribution and calculates the difference between the predicted and true mask distributions based on Wasserstein distance, thereby optimizing rotation-aware segmentation performance. More specifically, the RSReason model improves segmentation performance through its core Rotation Aware Optimization (RAO) module. The RAO module represents the target mask as a 2D Gaussian distribution and uses Wasserstein distance to measure the distribution difference between the predicted and true masks. In this way, the model can explicitly perceive and optimize the orientation and scale of the mask, thus significantly improving the segmentation accuracy for objects in any orientation, especially performing well when dealing with targets with complex orientation attributes. The evaluation dataset is input into the model to be evaluated, and the model is evaluated based on multi-dimensional metrics, calculating a benchmark evaluation score. During the evaluation process, the constructed RSReason dataset is input into the model to be evaluated. The model is comprehensively evaluated using multi-dimensional metrics, including gIoU, cIoU, and Acc@0.5. The model's performance on these metrics is calculated to obtain a benchmark evaluation score, which allows for a systematic measurement of the model's performance on remote sensing image inference and segmentation tasks.
[0053] Specifically, the geometric regularization injection mechanism approximates the mask as a 2D Gaussian ellipse, utilizes eigenvalue decomposition of the covariance matrix to derive the ellipse's principal axes and orientations, and converts geometric attributes into natural language descriptions, including descriptions of position, size, and orientation, as well as functional role inferences, generating inferential expressions. The formula for approximating the mask as a 2D Gaussian ellipse is:
[0054]
[0055] Where x, y represent the coordinates of each pixel in the mask in the Cartesian coordinate system; μ x ,μ yThese represent the center positions of the mask on the x-axis and y-axis, respectively; M(x,y) represents the function that generates the binary mask. These represent the variances of the mask on the x-axis and y-axis, respectively; σ xy Σ represents the covariance of the mask in the x and y directions; Σ represents the covariance matrix of the 2D Gaussian distribution into which the mask is converted; G(m) represents the 2D Gaussian distribution into which the mask is converted, where m represents the target binary mask; μ represents the mean of the 2D Gaussian distribution into which the mask is converted.
[0056] Specifically, the multi-dimensional evaluation metrics include gIoU, cIoU, and Acc@0.5, which are used to comprehensively measure the segmentation performance of the model on targets of different scales and orientations. The different scales include small, medium, and large objects, and the orientations include targets with obvious rotation and targets without obvious rotation.
[0057] Specifically, the overall architecture of the RSReason model includes a multimodal language module, a visual encoder, a mask decoder, and a RAO module. The multimodal language module jointly infers the image and text to generate a text sequence with embedded segmentation markers. The visual encoder extracts multi-scale features from the image. The mask decoder fuses language embeddings and visual features to generate a segmentation mask. The RAO module optimizes the orientation alignment of the mask. The RAO module of the RSReason model converts the mask into a 2D Gaussian distribution. Specifically, it constructs a covariance matrix by calculating the total mass, centroid, variance, and covariance of the mask, thereby obtaining the position, size, and orientation parameters of the mask, enhancing the model's ability to perceive rotating targets. The specific calculation method is as described above using the geometric regularization injection mechanism. The Wasserstein distance calculation formula used by the RAO module is:
[0058]
[0059] in, and represents the Gaussian distribution of the predicted and true masks, respectively; Tr represents the trace operation, which calculates the sum of the diagonal elements of the matrix.
[0060] Verification experiments were conducted on the embodiments of the present invention, and the experimental results are shown in Tables 1 and 2.
[0061] Table 1 - Evaluation scores of this invention and other models on the RSReason benchmark dataset:
[0062] Table 2 - Evaluation results of this invention and other models on the RRSIS-D dataset:
[0063]
[0064] As can be seen from the above, RSReason has better reasoning and segmentation capabilities.
[0065] join Figure 3 The diagram shown is a system structure diagram according to an embodiment of the present invention, including:
[0066] The data acquisition module 301 acquires image-text-mask pairs from the remote sensing image dataset, extracts the geometric attributes of the mask, generates an inferential description containing spatial, shape and functional cues through a geometric regularization injection mechanism, and constructs an evaluation dataset.
[0067] The model building module 302 constructs the RSReason model, inputs the evaluation dataset into the RSReason model, and trains and evaluates the model based on multi-dimensional indicators. During training, the RAO module is used to calculate the RAO loss function for each segmentation example, and the target mask is represented as a 2D Gaussian distribution. The difference between the predicted and the real mask distribution is calculated based on the Wasserstein distance. The model parameters are updated using this loss function to optimize the rotation-aware segmentation performance.
[0068] The image segmentation module 303 inputs the image to be segmented and the text describing the object to be segmented into the trained RSReason model for segmentation, and obtains the segmentation mask as the segmentation result.
[0069] The RSReason model includes:
[0070] A multimodal language module jointly infers image and text to generate text sequences with embedded segmentation tags;
[0071] A visual encoder extracts multi-scale features from an image;
[0072] The mask decoder fuses segmentation markers and multi-scale image features to generate a segmentation mask.
[0073] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A remote sensing reasoning segmentation method based on a visual language model, characterized in that, Includes the following steps: Image-text-mask pairs are obtained from remote sensing image datasets. Geometric attributes of the masks are extracted, and inferential descriptions containing spatial, shape, and functional cues are generated through a geometric regularization injection mechanism to construct an evaluation dataset. Construct an RSReason model by inputting the evaluation dataset into the RSReason model, and train and evaluate the model based on multi-dimensional metrics; During training, the RAO module is used to calculate the RAO loss function for each segmentation example. The target mask is represented as a 2D Gaussian distribution. The difference between the predicted and the true mask distribution is calculated based on the Wasserstein distance. The model parameters are updated using this loss function to optimize the rotation-aware segmentation performance. The image to be segmented and the text describing the object to be segmented are input into the trained RSReason model for segmentation, and the segmentation mask is obtained as the segmentation result. The RSReason model includes: A multimodal language module jointly infers image and text to generate text sequences with embedded segmentation tags; A visual encoder extracts multi-scale features from an image; The mask decoder fuses segmentation markers and multi-scale image features to generate a segmentation mask.
2. The remote sensing reasoning segmentation method based on a visual language model according to claim 1, characterized in that, The generation of inferential descriptions containing spatial, shape, and functional cues through the geometric regularization injection mechanism includes: The geometric regularization injection mechanism approximates the mask as a 2D Gaussian ellipse, uses eigenvalue decomposition of the covariance matrix to derive the principal axis and direction of the ellipse, and converts the geometric attributes into natural language descriptions, including descriptions of position, size and direction, as well as functional role inferences, to generate inferential expressions.
3. The remote sensing reasoning segmentation method based on a visual language model according to claim 2, characterized in that, The geometric regularization injection mechanism approximates the mask as a 2D Gaussian ellipse, as follows: Where x, y represent the coordinates of each pixel in the mask in the Cartesian coordinate system; μ x ,μ y These represent the center positions of the mask on the x-axis and y-axis, respectively; M(x,y) represents the function that generates the binary mask. These represent the variances of the mask on the x-axis and y-axis, respectively; σ xy Σ represents the covariance of the mask in the x and y directions; Σ represents the covariance matrix of the 2D Gaussian distribution into which the mask is converted; G(m) represents the 2D Gaussian distribution into which the mask is converted, where m represents the target binary mask; μ represents the mean of the 2D Gaussian distribution into which the mask is converted.
4. The remote sensing reasoning segmentation method based on a visual language model according to claim 1, characterized in that, The multi-dimensional metrics include gIoU, cIoU, and Acc@0.5, which are used to comprehensively measure the segmentation performance of the model on targets of different scales and orientations. Different scales include small, medium, and large objects, and orientations include targets with obvious rotation and targets without obvious rotation.
5. The remote sensing reasoning segmentation method based on a visual language model according to claim 1, characterized in that, The RAO module converts the mask into a 2D Gaussian distribution by calculating the total mass, centroid, variance, and covariance of the mask to construct a covariance matrix, thereby obtaining the position, size, and orientation parameters of the mask and enhancing the model's ability to perceive rotating targets.
6. The remote sensing reasoning segmentation method based on a visual language model according to claim 5, characterized in that, The difference between the predicted mask distribution based on Wasserstein distance calculation and the actual mask distribution is expressed as follows: in, and represents the Gaussian distribution of the predicted and true masks, respectively; Tr represents the trace operation, which calculates the sum of the diagonal elements of the matrix.
7. A remote sensing reasoning and segmentation system based on a visual language model, characterized in that, include: The data acquisition module obtains image-text-mask pairs from the remote sensing image dataset, extracts the geometric attributes of the masks, and generates inferential descriptions containing spatial, shape, and functional cues through a geometric regularization injection mechanism to construct the evaluation dataset. The model building module constructs the RSReason model, inputs the evaluation dataset into the RSReason model, and trains and evaluates the model based on multi-dimensional metrics. During training, the RAO module is used to calculate the RAO loss function for each segmentation example, represents the target mask as a 2D Gaussian distribution, calculates the difference between the predicted and the true mask distribution based on the Wasserstein distance, and uses this loss function to update the model parameters and optimize the rotation-aware segmentation performance. The image segmentation module takes the image to be segmented and the text describing the object to be segmented as input into the trained RSReason model for segmentation and obtains the segmentation mask as the segmentation result. The RSReason model includes: A multimodal language module jointly infers image and text to generate text sequences with embedded segmentation tags; A visual encoder extracts multi-scale features from an image; The mask decoder fuses segmentation markers and multi-scale image features to generate a segmentation mask.