Free text guided remote sensing image anaphora segmentation method and system
Through the regional relationship-driven graphic segmentation model, the problem of difficult aggregation of multi-objective expressions in remote sensing images is solved, and high-precision segmentation and stability improvement of complex remote sensing scenarios is achieved.
Patent Information
- Application Number
- CN202510823159.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-19
AI Technical Summary
The existing remote sensing image reference segmentation method has problems such as difficulty in aggregation of multi-objective expressions, easy to misjudgment of unobjective descriptions, and semantic fragmentation of multi-scale objects in complex and changeable remote sensing scenarios. The traditional method has insufficient generalization ability and robustness in open remote sensing scenarios.
The graphic and text segmentation model driven by region-relationship is adopted, including dynamic correlation vision encoder, pixel-level decoder, context-related text encoder, region-relationship modeling module and goal-oriented joint decoder, and multi-objective merging and target-free diagnosis are achieved through multi-scale perception, dynamic response enhancement, region-visual and region-language modeling interaction.
It significantly improves the joint identification and segmentation accuracy of multiple instance targets in complex remote sensing scenarios, enhances the accuracy, stability and generalization capabilities in open remote sensing scenarios, and improves the discrimination ability and stability under free expression of natural language.
Smart Images

Figure CN120340034A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image segmentation, and particularly to a free-text-guided remote sensing image referential segmentation method and system. Background Art
[0002] Remote sensing images refer to image data obtained by non-contact observation of the earth's surface using imaging sensors carried by satellites, aircraft, drones, etc. They have significant advantages such as wide coverage, high spatio-temporal resolution, and fast acquisition efficiency. With the continuous improvement of the performance of modern remote sensing sensors and the increase in observation frequencies, remote sensing images present higher spatial resolution and richer ground object information, making the types of ground objects contained in the images more complex and the morphological scale differences larger. Traditional image processing methods centered on visual features gradually expose problems in modern ground object recognition tasks, such as poor accuracy, robustness, and insufficient adaptability to complex scenes.
[0003] To solve these problems, in recent years, multi-modal modeling methods that integrate natural language and visual features have gradually been applied to remote sensing image analysis. The remote sensing image referential segmentation method that guides the model to identify regions with consistent semantics in the image through natural language descriptions provides a more flexible and controllable solution for remote sensing image segmentation and has become an important development direction for remote sensing intelligent analysis.
[0004] However, existing remote sensing image referential segmentation methods usually perform cross-modal alignment after encoding image and text features. These methods are mostly based on single-object and strong image-text pairing, relying on strict one-to-one semantic associations to achieve visual localization of specific targets by the model. However, in the face of complex and changing remote sensing scenes, this method still has significant limitations. First, there are many types of ground objects in remote sensing images, with large scale differences, and there are often multiple target objects with similar semantics or spatially adjacent. Traditional segmentation methods based on a single-object assumption are difficult to support "multi-object" or "regional group" expressions, easily leading to matching omissions or positioning deviations. Second, most existing methods rely on structured and templated text expressions for image-text alignment and are difficult to adapt to complex expression forms with free combinations of position, attribute, quantity, etc. in natural language, restricting flexible human-computer interaction capabilities. Finally, in the actual implementation process of remote sensing tasks, due to complex image backgrounds, blurred target boundaries, and strong interference, the generalization ability and robustness of existing models in open remote sensing scenes are insufficient, easily resulting in problems such as misidentification and false responses, affecting segmentation accuracy and system reliability. Summary of the Invention
[0005] To solve the above technical problems existing in the prior art, the present invention provides a free-text-guided remote sensing image referential segmentation method and system, and the technical solutions are as follows:
[0006] On the one hand, a free-text-guided remote sensing image referential segmentation method is provided, which includes:
[0007] S1. Collect and preprocess remote sensing image data;
[0008] S2. Perform instance-level object mask annotation and other processing on the preprocessed image data to obtain various labels, construct diverse natural language descriptions of the text, and build a data sample including images, texts, and various labels;
[0009] S3. Input the data sample into and train a region-relation-driven image-text segmentation model, where the region-relation-driven image-text segmentation model includes a dynamic association visual encoder, a pixel-level decoder, a context-associated text encoder, a region-relation modeling module, and a target-oriented joint decoder module;
[0010] The dynamic association visual encoder performs multi-scale perception and dynamic response enhancement on the input image data, and generates multi-scale visual features with spatial structure information ;
[0011] The pixel-level decoder performs pixel-level decoding on the multi-scale visual features and outputs image mask information including various category instances of the image ;
[0012] The context-associated text encoder performs semantic modeling on the input text, comprehensively extracts various key information included therein, and generates attribute-object information with context structure perception ability ;
[0013] The region-relation modeling module performs region-visual modeling interaction and region-language modeling interaction on the and respectively, gradually integrates various semantic information, improves the model's understanding and modeling ability of complex expressions, and obtains a region filter and region association features ;
[0014] The target-oriented joint decoder performs joint decoding on the , and , determines whether there is a true semantic match between the input text and the image, and performs multi-target merging and targetless diagnosis to achieve the multi-head prediction output of the model: target mask , region probability and target presence discrimination ;
[0015] S4. Use the trained region relationship-driven image-text segmentation model to segment the remote sensing image to be segmented.
[0016] Optionally, the S2 specifically includes:
[0017] Perform instance-level object mask annotation on the preprocessed image to obtain pixel mask labels ;
[0018] Perform spatial region division on the pixel mask, and compress the pixel-level annotation in each region into a value through average pooling, which is used to represent the probability that each region in the image includes a specific target category, as the region probability label , to reflect the spatial position and its distribution trend of the target in the image, and be used as a guiding signal in the subsequent training process to help the model focus on key regions;
[0019] Count whether each image includes a certain type of target. If the pixel annotation corresponding to the category in the image is not empty, mark this category as "1"; otherwise, mark it as "0" to obtain the target presence label , which is used to represent the presence of each target category in the image and is used as an auxiliary supervision signal in the subsequent training process;
[0020] Construct diverse natural language descriptions of this article. The text design covers six types of expression structures, including counting and ordinal expression, nested logical structure expression, multi-object attribute differentiation expression, complex relationship interaction expression, irrelevant object definition, and deceptive attribute design, comprehensively covering challenging description forms in remote sensing scenarios and enhancing the model's ability to understand and discriminate natural language;
[0021] Construct to obtain data samples including images , texts , pixel mask labels , region probability labels and target presence labels .
[0022] Optionally, the dynamic association visual encoder first performs preliminary projection and normalization on the image features through hidden space affine operations to obtain the original image encoding , improving the expression stability of the input;
[0023] Then, use multiple convolutional kernels with different receptive fields to perform multi-path parallel convolution operations on the image features to capture structural information and spatial context at different scales. The multi-scale features output by the multi-path parallel convolution are averaged point-by-point along the channel direction at each spatial position and fused into a unified feature, which not only retains scale diversity but also avoids interference caused by channel redundancy;
[0024] The unified features after fusion are fed into the Sigmoid activation function to generate pixel-level response weights, which are used to represent the saliency degree of pixels in each region in the global view. Multiply the pixel-level response weights with the multi-scale features point by point to achieve feature enhancement and suppression, and obtain the semantically enhanced features ;
[0025] Then, the is input into the visual gating unit, which combines the with the learnable transformation matrix, applies non-linear activation along the channel dimension, generates corresponding adjustment weights for each channel to constrain the , and finally outputs the multi-scale visual features . The formula is as follows:
[0026]
[0027] where represents the Sigmoid activation function, represents the learnable transformation matrix, represents element-wise multiplication for each channel.
[0028] Optionally, for the context-aware text encoder, first, map the input text to a vector representation through semantic embedding, introduce TextBlob for part-of-speech analysis, and extract key syntactic structures;
[0029] Then, the text vector and the key syntactic structures are respectively input into the text global feature extractor and the text local feature extractor constructed based on the BERT model to model long-range dependencies and local semantic details, comprehensively capture the target information and spatial relationships in the description, and obtain the local feature and the global feature ;
[0030] The and are fused through the gated weighted fusion module to achieve dynamic fusion of multi-granularity semantic features, including:
[0031] Concatenate the and , extract the fusion weights through the learnable linear transformation matrix , and normalize through the Sigmoid activation function to generate a weight gating factor, which is used to control the fusion ratio of the two-way features, and finally output the fused attribute-object information , realizing the adaptive combination of multi-granularity information at the semantic level. The formula is as follows:
[0032]
[0033] Among them represents the Sigmoid activation function, represents a learnable transformation matrix.
[0034] Optionally, the region relationship modeling module includes a region-vision cross-fusion RVI sub-module and a region-language cross-fusion RLI sub-module;
[0035] Among them, the RVI sub-module dynamically extracts the features of semantically related regions in the image through a region-level attention mechanism, including:
[0036] The original input image is segmented according to a fixed window size to obtain representative regions. The representative regions after segmentation are randomly initialized with learnable region query vectors , and combined with the learnable weight configuration to calculate the region-image association matrix , and the formula is as follows:
[0037]
[0038] Among them is the learnable parameter weight;
[0039] Based on the , through a linear perception layer for non-linear transformation and information compression, a region filter is obtained. The fuses the features of the region at different scales and positions, as well as the important information related to it in the global context, reflects the response intensity and content matching degree of each region in the image semantic structure, helps to distinguish the regions related to the target description from the irrelevant background, and is used for subsequent mask filtering;
[0040] According to the , the region visual perception features are extracted from the image, and the formula is as follows:
[0041]
[0042] Among them is the linear transformation weight;
[0043] The RLI sub-module models the dependence relationship between regions through an autocorrelation fitting mechanism to obtain region relationship perception features with global association modeling , and models the relationship between regions and language through a cross-correlation fitting mechanism to obtain region language perception features combining attribute-object information , then, the , and Perform weighted fusion to obtain region association features with spatial structure, context correlation, and language guidance capabilities ;
[0044] Among them, the autocorrelation fitting mechanism splits the by region division, performs linear transformation, calculates self-attention among regions, generates a similarity matrix between regions, applies the Softmax operation to normalize the similarity matrix, and then passes through a feed-forward neural network to obtain an attention weight matrix between regions. Then, the attention weight matrix between regions and are dot-multiplied and fused to obtain the ;
[0045] The cross-correlation fitting mechanism linearly transforms the and , calculates a response matrix of regions to language through cross-attention, normalizes the response matrix, and then captures non-linear features through a feed-forward neural network to obtain an attention matrix of region-language, representing the association strength between each region and each word in the language description. The attention matrix of region-language and are dot-multiplied and fused to obtain the .
[0046] Optionally, the target-oriented joint decoder jointly decodes the image instance pixel mask information and the region semantic features, effectively fusing region semantics, location perception, and language guidance information, including:
[0047] Receiving two types of outputs generated by the instance-oriented image mask information and region relationship modeling: region association features and region filters as inputs;
[0048] For the nth region, its corresponding region association feature outputs the confidence that this region includes the target through a linear perception layer. The region filter is point-multiplied with the image mask information to generate a region filtering mask that better conforms to the foreground response , representing its specific position and range in the image. The region filtering masks of all regions are weighted and aggregated according to their respective confidences to predict and output the overall target mask ;
[0049] Based on the confidence that each region includes the target, output the region probability as a representation of the target existence range;
[0050] By performing global average pooling on it, dimensionality reduction is carried out to a fixed dimension, and a linear perception layer is used to obtain the target existence discrimination of the model on whether the specified object exists in the image .
[0051] Optionally, the loss function of the region relationship-driven image-text segmentation model includes:
[0052] The predicted target mask The pixel-level segmentation loss calculated by comparing with the pixel label, which is used to supervise the consistency between the pixel-level target mask and the true label. The standard binary cross-entropy loss is adopted and denoted as :
[0053]
[0054] where represents the target existence probability of the th pixel predicted by the model, represents the true mask label of the th pixel;
[0055] The predicted region probability The region-level cross-entropy loss calculated by comparing with the region label , to evaluate whether the model can correctly identify which regions contain the target:
[0056]
[0057] where represents the probability that the model predicts whether the nth region contains the target, represents the true probability label of the nth region;
[0058] The predicted target existence discrimination The target discrimination classification loss calculated by comparing with the target existence label. The binary classification loss is adopted and denoted as :
[0059]
[0060] where represents the probability that the model predicts whether the entire image contains the target, represents the target existence label of the image;
[0061] The above three loss terms: , , , jointly constitute the final loss function, which is used to synchronously optimize the image region segmentation ability, the region target discrimination ability, and the image-text semantic consistency modeling ability.
[0062] On the other hand, a free-text-guided remote sensing image referential segmentation system is provided, and the system includes:
[0063] A collection and preprocessing module, configured to collect and preprocess remote sensing image data;
[0064] A construction module, configured to perform instance-level object mask annotation and other processing on the preprocessed image data to obtain various labels, construct diverse natural language description texts, and construct data samples including images, texts, and various labels;
[0065] A training module, configured to input the data samples and train a region-relation-driven image-text segmentation model, where the region-relation-driven image-text segmentation model includes a dynamic association visual encoder, a pixel-level decoder, a context-associated text encoder, a region-relation modeling module, and a target-oriented joint decoder module;
[0066] The dynamic association visual encoder performs multi-scale perception and dynamic response enhancement on the input image data to generate multi-scale visual features with spatial structure information ;
[0067] The pixel-level decoder performs pixel-level decoding on the multi-scale visual features and outputs image mask information including various category instances of the image ;
[0068] The context-associated text encoder performs semantic modeling on the input text, comprehensively extracts various key information included therein, and generates attribute-object information with context structure perception ability ;
[0069] The region-relation modeling module performs region-visual modeling interaction and region-language modeling interaction on the and respectively, gradually integrates various semantic information, improves the model's understanding and modeling ability of complex expressions, and obtains a region filter and region association features ;
[0070] The target-oriented joint decoder performs joint decoding on the 、 and , determines whether there is a true semantic match between the input text and the image, and performs multi-target merging and targetless diagnosis to achieve the multi-head prediction output of the model: target mask 、region probability and target existence discrimination ;
[0071] A segmentation module, which is used to segment the remote sensing image to be segmented by using the region relationship-driven image-text segmentation model completed in training.
[0072] On the other hand, an electronic device is provided. The electronic device includes a processor and a memory. At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the above-mentioned free-text-guided remote sensing image referential segmentation method.
[0073] On the other hand, a computer-readable storage medium is provided. At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the above-mentioned free-text-guided remote sensing image referential segmentation method.
[0074] The beneficial effects brought by the technical solution provided by the present invention at least include:
[0075] 1) The region relationship modeling mechanism designed by the present invention realizes the context interaction modeling between multiple targets by constructing an explicit semantic association graph of regions in the image, supports the overall parsing and positioning of general and collective descriptions in natural language, and significantly improves the joint recognition and segmentation accuracy of multi-instance targets in complex remote sensing scenarios.
[0076] 2) The dynamic association visual encoder and context association text encoder designed by the present invention can flexibly adjust the region perception range according to the input semantics, and fuse multi-scale context semantic features to realize the accurate modeling and positioning of targets at different scales, significantly enhancing the accuracy, stability and generalization ability in open remote sensing scenarios.
[0077] 3) The target-oriented joint decoder designed by the present invention introduces a region-level cross-modal matching and existence judgment mechanism, which can effectively identify whether there is an actual visual correspondence in the input text, avoids incorrect responses to irrelevant descriptions, and greatly improves the discrimination ability and stability under free natural language expression.
[0078] 4) The present invention adopts a multi-task joint training strategy, weights and optimizes the losses of three tasks: target mask, region target probability and target existence discrimination, to ensure that the model achieves high accuracy in multiple tasks. This optimization strategy can improve the accuracy of target positioning, segmentation and discrimination while maintaining the stability of the system, making the model perform excellently in various complex remote sensing image analysis tasks. Description of the Drawings
[0079] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0080] Figure 1 It is a flowchart of a free-text-guided remote sensing image anaphoric segmentation method provided by an embodiment of the present invention;
[0081] Figure 2 It is a general block diagram of a free-text-guided remote sensing image anaphoric segmentation method provided by an embodiment of the present invention;
[0082] Figure 3 It is a schematic diagram of data sample construction provided by an embodiment of the present invention;
[0083] Figure 4 It is a block diagram of the region relationship-driven graph-text segmentation model structure provided by an embodiment of the present invention;
[0084] Figure 5 It is a block diagram of the dynamic association visual encoder structure provided by an embodiment of the present invention;
[0085] Figure 6 It is a block diagram of the context association text encoder structure provided by an embodiment of the present invention;
[0086] Figure 7 It is a block diagram of the region relationship modeling module structure provided by an embodiment of the present invention;
[0087] Figure 8 It is a schematic diagram of the self-correlation fitting mechanism provided by an embodiment of the present invention;
[0088] Figure 9 It is a schematic diagram of the cross-correlation fitting mechanism provided by an embodiment of the present invention;
[0089] Figure 10 It is a block diagram of the target-oriented joint decoder structure provided by an embodiment of the present invention;
[0090] Figure 11 It is a block diagram of a free-text-guided remote sensing image anaphoric segmentation system provided by an embodiment of the present invention;
[0091] Figure 12 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Specific embodiments
[0092] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the drawings and specific embodiments.
[0093] An embodiment of the present invention provides a free-text-guided remote sensing image referential segmentation method, which can be implemented by an electronic device, and the electronic device can be a terminal or a server. Figure 1 The flowchart of the method is shown as follows, Figure 2 The overall block diagram of the method is shown as follows, aiming to solve problems such as difficult aggregation of multi-object expressions, easy misjudgment of objectless descriptions, and semantic fragmentation of multi-scale objects in remote sensing image segmentation. The specific processing flow may include the following steps:
[0094] S1. Collect and preprocess remote sensing image data;
[0095] In an embodiment of the present invention, a high-resolution visible light sensor is carried by a drone to conduct low-altitude aerial photography of typical areas such as urban buildings, rural roads, farmlands, water bodies, etc., and collect remote sensing image data with clear structures and rich semantics. The images are original color images, with a unified size of 4000×3000 pixels and a resolution better than 0.2 meters / pixel, which can accurately present the outlines and positional relationships of common ground objects. During the collection process, the flight height and angle are controlled to ensure that the images are unobstructed and have small distortions, meeting the requirements of subsequent annotation and text construction.
[0096] After the collection of remote sensing image data is completed, in an embodiment of the present invention, the original images are preprocessed, including unifying format specifications and size cropping, and standardizing them to 512×512 pixels to meet the model input requirements.
[0097] S2. Perform instance-level object mask annotation and other processing on the preprocessed image data to obtain various labels, construct diverse natural language description texts, and build data samples including images, texts, and various labels;
[0098] Optionally, as Figure 3 shown, the S2 specifically includes:
[0099] Perform instance-level object mask annotation on the preprocessed image (accurately annotate the categories and boundaries of ground object targets such as buildings, roads, water bodies, vegetation, vehicles, etc. to ensure clear semantics and complete structures) to obtain pixel mask labels ;
[0100] Perform spatial region division on the pixel mask, and compress the pixel-level annotation in each region into a value through average pooling, which is used to represent the probability that each region in the image includes a specific target category, as the region probability label , to reflect the spatial position and its distribution trend of the target in the image, and be used as a guiding signal in the subsequent training process to help the model focus on key regions;
[0101] Statistically determine whether each image contains a certain type of target. If the pixel annotation of the corresponding category in the image is not empty, mark this category as "1"; otherwise, mark it as "0" to obtain the target presence label , which is used to represent the presence of each target category in the image and serves as an auxiliary supervision signal during the subsequent training process;
[0102] Construct diverse natural language descriptions of this article (free text). The text design covers six types of expression structures, including counting and ordinal expressions (used for target quantity and sorting and positioning), nested logical structure expressions (reflecting relationships such as parallelism and exclusion between targets), multi-target attribute differentiation expressions (describing targets with multiple attribute similarities and differences), complex relationship interaction expressions (involving spatial or semantic dependencies between targets), irrelevant object definitions (constructing expressions that are partially relevant but have no corresponding targets), and deceptive attribute designs (introducing interfering descriptions to enhance discrimination robustness), comprehensively covering challenging description forms in remote sensing scenarios and enhancing the model's understanding and discrimination ability of natural language;
[0103] Construct a data sample including the image , text , pixel mask label , region probability label and target presence label .
[0104] In addition, after the data sample construction is completed in the embodiments of the present invention, the data is divided into a training set and a validation set according to a ratio of 4:1, and different types of text-image pairs are reasonably organized, including three types of samples: one-to-one, one-to-many, and one-to-zero. For example, in the one-to-one type of sample, the text description is "the red-roofed building in the lower right corner", which only corresponds to a single ground object target in the image; in the one-to-many type of sample, the description is "the white cars on both sides of the road", which involves multiple semantically related instances in the image; in the one-to-zero type of sample, the description is "the pedestrian wearing a blue dress in the picture", while there is no corresponding target in the image. During the division process, the distribution of various types of samples in the training set and the validation set is kept balanced to enhance the generalization ability and evaluation stability of the model. Moreover, a background random occlusion strategy is introduced. During the training stage, non-target regions in the image are randomly selected, and square occlusions are performed in a fixed-size sliding window manner and replaced with mean filling or noise perturbation, thereby suppressing background interference and enhancing the model's semantic focus and discrimination ability for key target regions.
[0105] S3. Input the data sample into and train a region relationship-driven text-image segmentation model. The region relationship-driven text-image segmentation model includes a dynamic association visual encoder, a pixel-level decoder, a context association text encoder, a region relationship modeling module, and a target-oriented joint decoder module, as Figure 4 shown;
[0106] The dynamic association visual encoder performs multi-scale perception and dynamic response enhancement on the input image data to generate multi-scale visual features with spatial structure information ;
[0107] The pixel-level decoder performs pixel-level decoding on the multi-scale visual features and outputs image mask information including various category instances of the image ;
[0108] The context association text encoder performs semantic modeling on the input text, comprehensively extracts various key information included therein, and generates attribute-object information with context structure perception ability ;
[0109] The region relationship modeling module performs region-vision modeling interaction and region-language modeling interaction on the and respectively, gradually integrates various semantic information, improves the model's understanding and modeling ability of complex expressions, and obtains a region filter and region association features ;
[0110] The target-oriented joint decoder performs joint decoding on the 、 and to determine whether there is a true semantic match between the input text and the image, and performs multi-target merging and targetless diagnosis to achieve the multi-head prediction output of the model: target mask 、region probability and target existence discrimination ;
[0111] Optionally, as shown in Figure 5 , the dynamic association visual encoder first performs preliminary projection and normalization on the image features through latent space affine operations to obtain the original image encoding , improving the expression stability of the input;
[0112] Then, convolution kernels with multiple different receptive fields (including 3×3, 5×5, and 7×7) are used to perform multi-path parallel convolution operations on the image features to capture the structural information and spatial context at different scales. The multi-scale features output by the multi-path parallel convolution are averaged point by point along the channel direction at each spatial position and fused into a unified feature, which not only retains scale diversity but also avoids interference caused by channel redundancy;
[0113] The fused unified features are fed into the Sigmoid activation function to generate pixel-level response weights, which are used to represent the saliency degree of pixels in each region in the global view. The pixel-level response weights are multiplied element-wise with the multi-scale features to achieve feature enhancement and suppression, obtaining semantically enhanced features. ;
[0114] Then the is input into the visual gating unit, which combines the with a learnable transformation matrix, applies non-linear activation along the channel dimension, generates corresponding adjustment weights for each channel to constrain the , and finally outputs multi-scale visual features . The formula is as follows:
[0115]
[0116] where represents the Sigmoid activation function, represents the learnable transformation matrix, represents element-wise multiplication for each channel.
[0117] The visual gating unit of the embodiment of the present invention can dynamically regulate the channel activation distribution of features according to the guiding information, strengthen the feature dimensions consistent with the semantics of the language description, thereby effectively improving the cross-modal alignment ability and discrimination performance.
[0118] Optionally, as Figure 6 shown, the context-associated text encoder, first, maps the input text into a vector representation through semantic embedding, and introduces TextBlob for part-of-speech analysis to extract key syntactic structures;
[0119] Then the text vector and the key syntactic structures are respectively passed through a text global feature extractor and a text local feature extractor constructed based on the BERT model to model long-range dependencies and local semantic details, comprehensively capture the target information and spatial relationships in the description, and obtain local features and global features ;
[0120] The and are fused dynamically through a gated weighted fusion module, including:
[0121] The and are concatenated, and through a learnable linear transformation matrix Extract the fusion weights, normalize them via the Sigmoid activation function to generate a weight gating factor for controlling the fusion ratio of the two-way features, and finally output the fused attribute-object information. , achieving the adaptive combination of multi-granularity information at the semantic level, which is expressed by the following formula:
[0122]
[0123] where represents the Sigmoid activation function, represents a learnable transformation matrix.
[0124] Optionally, as shown in Figure 7 , the region relationship modeling module includes a Region-Vision Integration (RVI) sub-module and a Region-Language Integration (RLI) sub-module;
[0125] Among them, the RVI sub-module dynamically extracts the features of semantically related regions in the image through a region-level attention mechanism, including:
[0126] Slice the original input image according to a fixed window size to obtain representative regions, randomly initialize the learnable region query vector for the sliced representative regions, and combine it with the to calculate the region-image association matrix , and the formula is as follows:
[0127]
[0128] where is the learnable parameter weight;
[0129] Based on the , perform non-linear transformation and information compression through a linear perception layer to obtain the region filter , and the integrates the features of the region at different scales and positions, as well as the important information related to it in the global context, reflecting the response intensity and content matching degree of each region in the image semantic structure, helping to distinguish the regions related to the target description from the irrelevant background for subsequent mask filtering;
[0130] According to the , extract the region visual perception features from the image, and the formula is as follows:
[0131]
[0132] wherein is the weight of the linear transformation;
[0133] The above mechanism enables each regional feature to be dynamically aggregated from the whole image, with higher flexibility and adaptability.
[0134] The RLI sub-module models the dependencies between regions through the self-correlation fitting mechanism to obtain the region relation perception features with global correlation modeling and models the relationship between regions and language through the cross-correlation fitting mechanism to obtain the region language perception features combining attribute-object information Then, the , and are weighted and fused to obtain the region association features with spatial structure, context association and language guidance capabilities ;
[0135] Among them, the self-correlation fitting mechanism, as Figure 8 shown, splits the by region, performs self-attention calculation between regions after linear transformation, generates a similarity matrix between regions, applies the Softmax operation to normalize the similarity matrix, and then passes through a feed-forward neural network to obtain the attention weight matrix between regions. Then, the attention weight matrix between regions and are dot-multiplied and fused to obtain the ;
[0136] This mechanism effectively captures the semantic associations between regions in remote sensing images with multiple targets and complex backgrounds, helping to overcome the semantic ambiguity problem.
[0137] The cross-correlation fitting mechanism, as Figure 9 shown, after the and undergo their respective linear transformations, generate a response matrix of regions to language through cross-attention calculation. After the response matrix is normalized, the feed-forward neural network captures non-linear features to obtain the region-language attention matrix, representing the association strength between regions and each word in the language description. The region-language attention matrix and are dot-multiplied and fused to obtain the .
[0138] This mechanism enhances the semantic alignment ability between the language description and the image content, thus improving the accuracy of target localization and segmentation.
[0139] Optionally, as Figure 10As shown, the target-oriented joint decoder jointly decodes the image instance pixel mask information and the regional semantic features, effectively integrating regional semantics, location awareness, and language guidance information, including:
[0140] Receiving instance-oriented image mask information And two types of outputs generated by regional relationship modeling: regional association features And regional filters As inputs;
[0141] For the nth region, its corresponding regional association feature Outputs the confidence that this region contains the target through a linear perception layer. The regional filter Is multiplied element-wise with the image mask information To generate a regional filtering mask that better conforms to the foreground response , representing its specific position and range in the image. The regional filtering masks of all regions Are weighted and aggregated according to their respective confidences to predict and output the overall target mask ;
[0142] Based on the confidence that each region contains the target, outputs the regional probability As a representation of the target existence range;
[0143] By performing global average pooling on , reducing its dimension to a fixed dimension, and using a linear perception layer to obtain the target existence discrimination of whether the specified object exists in the image .
[0144] This joint decoding mechanism effectively integrates regional semantics, location awareness, and language guidance information, has strong spatial reasoning and expression matching capabilities, and significantly improves the segmentation accuracy and robustness in multi-object and multi-relationship scenarios of remote sensing images.
[0145] Optionally, the loss function of the regional relationship-driven graph-text segmentation model includes:
[0146] The predicted target mask And the pixel-level segmentation loss calculated by comparing with the pixel labels, which is used to supervise the consistency between the pixel-level target mask and the true label. The standard binary cross-entropy loss is used, denoted as :
[0147]
[0148] Where Represents the target existence probability of the th pixel predicted by the model, Represents the The true mask label of a pixel;
[0149] The predicted region probability The region-level cross-entropy loss calculated by comparing with the region label , to evaluate whether the model can correctly identify which regions contain the target:
[0150]
[0151] where represents the probability that the model predicts whether the nth region contains the target, represents the true probability label of the nth region;
[0152] The predicted target existence discrimination The target discrimination classification loss calculated by comparing with the target existence label, using the binary classification loss, denoted as :
[0153]
[0154] where represents the probability that the model predicts whether the entire image contains the target, represents the target existence label of the image;
[0155] The above three loss terms: , , , jointly constitute the final loss function, used to synchronously optimize the image region segmentation ability, region target discrimination ability, and the ability to model the consistency of image-text semantics.
[0156] S4. Use the trained region relationship-driven image-text segmentation model to segment the remote sensing image to be segmented.
[0157] As Figure 11 shown, the embodiment of the present invention also provides a free text-guided remote sensing image referential segmentation system, and the system includes:
[0158] The collection and preprocessing module 1110, used to collect and preprocess remote sensing image data;
[0159] The construction module 1120, used to perform instance-level target mask annotation and other processing on the preprocessed image data to obtain various labels, construct diverse natural language descriptions of this article, and construct data samples including images, texts, and various labels;
[0160] A training module 1130, configured to input the data samples and train a region-relation-driven image-text segmentation model, where the region-relation-driven image-text segmentation model includes a dynamic association visual encoder, a pixel-level decoder, a context-associated text encoder, a region-relation modeling module, and a target-oriented joint decoder module;
[0161] The dynamic association visual encoder performs multi-scale perception and dynamic response enhancement on the input image data to generate multi-scale visual features with spatial structure information ;
[0162] The pixel-level decoder performs pixel-level decoding on the multi-scale visual features to output image mask information including various category instances of the image ;
[0163] The context-associated text encoder performs semantic modeling on the input text, comprehensively extracts various key information included therein, and generates attribute-object information with context structure perception ability ;
[0164] The region-relation modeling module performs region-vision modeling interaction and region-language modeling interaction on the and respectively, gradually integrates various semantic information, improves the model's understanding and modeling ability for complex expressions, and obtains a region filter and region association features ;
[0165] The target-oriented joint decoder performs joint decoding on the , and to determine whether there is a true semantic match between the input text and the image, and perform multi-target merging and target-free diagnosis to achieve the multi-head prediction output of the model: a target mask , a region probability and a target existence determination ;
[0166] A segmentation module 1140, configured to use the trained region-relation-driven image-text segmentation model to segment the remotely sensed image to be segmented.
[0167] A free-text-guided remotely sensed image referential segmentation system provided by an embodiment of the present invention has a functional structure corresponding to a free-text-guided remotely sensed image referential segmentation method provided by an embodiment of the present invention, and will not be elaborated herein.
[0168] Figure 12It is a schematic structural diagram of an electronic device 1200 provided by an embodiment of the present invention. The electronic device 1200 may vary greatly due to different configurations or performances, and may include one or more processors (central processing units, CPUs) 1201 and one or more memories 1202. Among them, at least one instruction is stored in the memory 1202, and the at least one instruction is loaded and executed by the processor 1201 to implement the steps of the above-mentioned free-text-guided remote sensing image referential segmentation method.
[0169] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions, and the above instructions can be executed by a processor in a terminal to complete the above-mentioned free-text-guided remote sensing image referential segmentation method. For example, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0170] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a magnetic disk, or an optical disc, etc.
[0171] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A free-text-guided remote sensing image referential segmentation method, characterized in that The method includes: S1. Collect and preprocess remote sensing image data; S2. Perform instance-level object mask annotation and other processing on the preprocessed image data to obtain various labels, construct diverse natural language descriptions of the text, and build a data sample including images, text, and various labels; S3. Input the data sample into and train a region-relation-driven image-text segmentation model, where the region-relation-driven image-text segmentation model includes a dynamic association visual encoder, a pixel-level decoder, a context-associated text encoder, a region-relation modeling module, and a target-oriented joint decoder module; The dynamic association visual encoder performs multi-scale perception and dynamic response enhancement on the input image data to generate multi-scale visual features with spatial structure information. ; The pixel-level decoder decodes the multi-scale visual features at the pixel level and outputs image mask information including various category instances of the image ; The context-related text encoder performs semantic modeling on the input text, comprehensively extracts various key information included therein, and generates attribute-object information with context structure awareness ability ; The region relationship modeling module performs region-visual modeling interaction and region-language modeling interaction on the and respectively, gradually integrating various types of semantic information, enhancing the model's understanding and modeling capabilities for complex expressions, and obtaining a region filter and region association features ; The target-oriented joint decoder jointly decodes the , and to determine whether there is a true semantic match between the input text and the image, and performs multi-target merging and targetless diagnosis to achieve the multi-head prediction output of the model: target mask , region probability and target presence discrimination ; S4. Use the trained region-relation-driven image-text segmentation model to segment the remote sensing image to be segmented.
2. The method according to claim 1, wherein The S2 specifically includes: Perform instance-level object mask annotation on the preprocessed image to obtain pixel mask labels ; Perform spatial region division on the pixel mask, and compress the pixel-level annotations in each region into a value through average pooling, which is used to represent the probability that each region in the image includes a specific target category, serving as the region probability label , so as to reflect the spatial position and its distribution trend of the target in the image, and be used as a guiding signal in the subsequent training process to help the model focus on the key regions; Statistically determine whether each image contains a certain type of target. If the pixel annotation for the corresponding category in the image is not empty, mark this category as "1"; otherwise, mark it as "0" to obtain the target presence label , which is used to represent the presence of each target category in the image and serves as an auxiliary supervision signal during the subsequent training process; Construct diverse natural language descriptions of the text. The text design covers six types of expression structures, including counting and ordinal expression, nested logical structure expression, multi-object attribute differentiation expression, complex relationship interaction expression, irrelevant object definition, and deceptive attribute design, comprehensively covering challenging description forms in remote sensing scenarios and enhancing the model's understanding and discrimination ability of natural language; Construct a data sample including an image , text , a pixel mask label , a region probability label , and a target presence label .
3. The method according to claim 1, wherein The dynamic association visual encoder first performs preliminary projection and normalization on the image features through latent space affine operations to obtain the original image encoding , improving the expression stability of the input; Then, use convolutional kernels with multiple different receptive fields to perform multi-path parallel convolutional operations on the image features to capture structural information and spatial context at different scales. The multi-scale features output by the multi-path parallel convolution are averaged point-by-point along the channel direction at each spatial position and fused into a unified feature, which not only retains scale diversity but also avoids interference caused by channel redundancy; The fused unified features are fed into the Sigmoid activation function to generate pixel-level response weights, which are used to represent the saliency degree of pixels in each region in the global view. Multiply the pixel-level response weights and the multi-scale features point by point to achieve feature enhancement and suppression, and obtain semantically enhanced features ; Then, the is input into the visual gating unit, which combines the with a learnable transformation matrix, applies non-linear activation along the channel dimension, generates corresponding adjustment weights for each channel to constrain the , and finally outputs multi-scale visual features , which is expressed by the formula as follows: ; wherein represents the Sigmoid activation function, represents a learnable transformation matrix, represents element-wise multiplication for each channel.
4. The method according to claim 1, wherein For the context-associated text encoder, first, map the input text to a vector representation through semantic embedding, and introduce TextBlob for part-of-speech analysis to extract key syntactic structures; Then, the text vector and the key grammatical structures are respectively input into the text global feature extractor and the text local feature extractor constructed based on the BERT model to model long-distance dependencies and local semantic details, comprehensively capture the target information and spatial relationships in the description, and obtain local features and global features ; The said and Through the gated weighted fusion module, the dynamic fusion of multi-granularity semantic features is realized, including: Concatenate the and , extract the fusion weights through a learnable linear transformation matrix , normalize them via the Sigmoid activation function to generate a weight gating factor for controlling the fusion ratio of the two-way features, and finally output the fused attribute-object information , achieving an adaptive combination of multi-granularity information at the semantic level, and the formula is expressed as follows: ; wherein represents the Sigmoid activation function, represents a learnable transformation matrix.
5. The method according to claim 1, wherein The region-relation modeling module includes a region-vision cross-fusion RVI sub-module and a region-language cross-fusion RLI sub-module; Among them, the RVI sub-module dynamically extracts the features of semantically related regions in the image through a region-level attention mechanism, including: The original input image is segmented according to a fixed window size to obtain representative regions. The randomly initialized learnable region query vectors are used for the segmented representative regions and combined with the learnable weight configuration to calculate the region-image association matrix The formula is as follows: ; wherein is the learnable parameter weight; Based on the above , nonlinear transformation and information compression are performed through a linear perception layer to obtain a regional filter , the integrates the features of the region at different scales and positions, as well as the important information related to it in the global context, reflects the response intensity and content matching degree of each region in the image semantic structure, helps to distinguish the regions related to the target description from the irrelevant background, and is used for subsequent mask filtering; According to the said , extract the regional visual perception features from the image , and the formula is as follows: ; wherein is the weight of the linear transformation; The RLI sub-module models the dependencies between regions through an autocorrelation fitting mechanism to obtain region relation-aware features with global correlation modeling , models the relationship between regions and language through a cross-correlation fitting mechanism to obtain region language-aware features that combine attribute-object information , then, the , and are weighted and fused to obtain region association features with spatial structure, context correlation, and language guidance capabilities ; Among them, the autocorrelation fitting mechanism is split according to regional division. After linear transformation, self-attention calculation is performed between regions to generate a similarity matrix between regions. The Softmax operation is applied to normalize the similarity matrix, and then through a feed-forward neural network, an inter-region attention weight matrix is obtained. Then, the inter-region attention weight matrix and are dot-product fused to obtain the ; The cross-correlation fitting mechanism will and After passing through their respective linear transformations, the response matrix of the region to the language is generated through cross-attention calculation. After the response matrix is normalized, the feed-forward neural network captures the non-linear features to obtain the region-language attention matrix, which represents the association strength between the region and each word in the language description. The region-language attention matrix and are dot-multiplied and fused to obtain the .
6. The method according to claim 1, wherein The target-oriented joint decoder jointly decodes the image instance pixel mask information and the region semantic features, effectively fusing region semantics, location perception, and language guidance information, including: Receive instance-oriented image mask information Two types of outputs generated by region relationship modeling: region association features and region filters as inputs; For the nth region, its corresponding region correlation feature Output the confidence that the target is included in this region through a linear perception layer, and the region filter Perform point-by-point multiplication with the image mask information to generate a region filtering mask that better conforms to the foreground response , indicating its specific position and range in the image. The region filtering masks of all regions are weighted and aggregated according to their respective confidences to predict and output the overall target mask ; Output the region probability based on the confidence of each region including the target As a representation of the target existence range; By performing global average pooling on it to reduce its dimension to a fixed dimension, and using a linear perception layer to obtain the target existence discrimination of whether the specified object exists in the image by the model .
7. The method according to claim 1, wherein The loss function of the region-relation-driven image-text segmentation model includes: Predicted target mask The pixel-level segmentation loss calculated by comparing with the pixel labels is used to supervise the consistency between the pixel-level target mask and the ground truth labels. The standard binary cross-entropy loss is used, denoted as : ; Among them represents the target existence probability of the th pixel predicted by the model, represents the true mask label of the th pixel; Predicted region probability Region-level cross-entropy loss calculated by comparing with region labels , to evaluate whether the model can correctly identify which regions contain the target: ; wherein represents the probability that the model predicts whether the nth region includes the target, represents the true probability label of the nth region; Predicted target existence discrimination The target discrimination classification loss calculated by comparing with the target existence label adopts a binary classification loss, denoted as : ; wherein represents the probability that the model predicts whether the entire image includes the target, represents the target presence label of the image; The above three loss terms: , , , jointly constitute the final loss function, which is used to synchronously optimize the image region segmentation ability, region target discrimination ability, and graphic-text semantic consistency modeling ability.
8. A free-text-guided remote sensing image referential segmentation system, characterized in that, The system includes: A collection and preprocessing module for collecting and preprocessing remote sensing image data; A construction module for performing instance-level object mask annotation and other processing on the preprocessed image data to obtain various labels, constructing diverse natural language descriptions of the text, and building a data sample including images, text, and various labels; A training module for inputting the data sample into and training a region-relation-driven image-text segmentation model, where the region-relation-driven image-text segmentation model includes a dynamic association visual encoder, a pixel-level decoder, a context-associated text encoder, a region-relation modeling module, and a target-oriented joint decoder module; The dynamic association visual encoder performs multi-scale perception and dynamic response enhancement on the input image data to generate multi-scale visual features with spatial structure information ; The pixel-level decoder decodes the multi-scale visual features at the pixel level and outputs image mask information including various category instances of the image ; The context-related text encoder performs semantic modeling on the input text, comprehensively extracts various key information included therein, and generates attribute-object information with context structure awareness ability ; The region relationship modeling module performs region-visual modeling interaction and region-language modeling interaction on the and respectively, gradually integrating various types of semantic information to enhance the model's understanding and modeling capabilities for complex expressions, and obtaining a region filter and region association features ; The target-oriented joint decoder jointly decodes the , and to determine whether there is a true semantic match between the input text and the image, and performs multi-target merging and targetless diagnosis to achieve the multi-head prediction output of the model: target mask , region probability and target presence discrimination ; A segmentation module for using the trained region-relation-driven image-text segmentation model to segment the remote sensing image to be segmented.
9. An electronic device, the electronic device includes a processor and a memory, and at least one instruction is stored in the memory, characterized in that, The at least one instruction is loaded and executed by the processor to implement the free-text-guided remote sensing image co-reference segmentation method according to any one of claims 1-7.
10. A computer-readable storage medium storing at least one instruction, characterized in that, The at least one instruction is loaded and executed by the processor to implement the free-text-guided remote sensing image co-reference segmentation method according to any one of claims 1-7.
Citation Information
Patent Citations
Visual localization and anaphora segmentation method, system and device based on mask anaphora modeling and storage medium
CN118734091A
Medical visual question and answer method and system based on multi-task modeling
CN119202334A
Auricle anaphora segmentation method and system
CN119579905A
Image anaphora segmentation method based on autoregression vertex generation and language structure guidance
CN119850952A
Remote sensing visual positioning method based on multi-scale progressive reasoning
CN120107354A
Cited By
Wound repair effect prediction method based on generative adversarial network
CN120598821A
Remote sensing scene graph guided semantic information reasoning method and device, equipment and medium
CN120599616A
Text-guided human ear three-dimensional point cloud anaphora segmentation method and system
CN120853240A
Training reasoning method based on voice-text-image multi-mode contrast learning
CN121562829A
Semantic segmentation method and device for remote sensing image
CN121582592A