A free text-guided remote sensing image reference segmentation method and system

Through the regional relationship-driven graphic and text segmentation model, the problem of difficulty in aggregation of multi-objective expressions in remote sensing image reference segmentation is solved, and the problem of insufficient robustness in open scenarios is achieved, and high-precision segmentation and discrimination of complex remote sensing scenarios is achieved.

CN120340034BActive Publication Date: 2025-08-26UNIV OF SCI & TECH BEIJING
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510823159.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-08-26
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

Existing remote sensing images refer to segmentation methods that are difficult to support multi-objective or regional group expression when facing complex and changeable remote sensing scenarios, and their generalization ability and robustness in open remote sensing scenarios are insufficient, making them prone to misidentification and false responses.

Method used

The free text-guided remote sensing image reference segmentation method is used to drive the graphic and text segmentation model through regional relationships, including dynamic correlation vision encoder, pixel-level decoder, context-related text encoder, region-related modeling module and goal-oriented joint decoder, to build multi-scale visual features and context structure perception capabilities to achieve multi-objective merging and target-free diagnosis.

Benefits of technology

It significantly improves the joint identification and segmentation accuracy of multiple instance targets in complex remote sensing scenarios, enhances the accuracy, stability and generalization capabilities in open remote sensing scenarios, and improves the discrimination ability and stability under free expression of natural language.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340034B_ABST
    Figure CN120340034B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for remote sensing image referent segmentation guided by free text. The method comprises: constructing data samples including images, text, and various labels, inputting and training a region-relationship-driven image-text segmentation model. The model comprises a dynamic association visual encoder that performs multi-scale perception and dynamic response enhancement on images to generate multi-scale visual features #imgabs0#; a pixel-level decoder that performs pixel-level decoding on #imgabs1# and outputs image mask information #imgabs2#; a context-related text encoder that performs semantic modeling on text to generate attribute-object information #imgabs3#; a region-relationship modeling module that performs interactive region-vision and region-language modeling on #imgabs4# and #imgabs5#, respectively, to obtain region filters #imgabs6# and region-relation features #imgabs7#; and a target-oriented joint decoder that performs joint decoding on #imgabs8#, #imgabs9#, and #imgabs10# to achieve multi-head prediction output of the model. The present invention can segment remote sensing images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image segmentation, and in particular to a free text-guided remote sensing image reference segmentation method and system. Background Art

[0002] Remote sensing imagery refers to image data obtained through non-contact observation of the Earth's surface using imaging sensors mounted on satellites, aircraft, drones, and other vehicles. These images offer significant advantages, including wide coverage, high spatiotemporal resolution, and rapid acquisition efficiency. With the continuous improvement in the performance of modern remote sensing sensors and the increase in observation frequency, remote sensing images exhibit higher spatial resolution and richer ground object information, resulting in more complex types of ground objects and greater differences in morphology and scale. Traditional image processing methods, centered around visual features, are gradually exposing problems in modern ground object recognition tasks, such as poor accuracy, poor robustness, and insufficient adaptability to complex scenarios.

[0003] In order to solve these problems, in recent years, multimodal modeling methods that integrate natural language and visual features have gradually been applied to remote sensing image analysis. The remote sensing image reference segmentation method, which guides the model to identify semantically consistent areas in the image through natural language description, provides a more flexible and controllable solution for remote sensing image segmentation, and has become an important development direction of remote sensing intelligent analysis.

[0004] However, existing methods for referential segmentation in remote sensing images typically encode image and text features and then perform cross-modal alignment. These methods are often based on single-target, strong image-text pairing and rely on strict one-to-one semantic associations to achieve visual localization of specific targets. However, these methods still have significant limitations in the complex and ever-changing remote sensing scenes. First, remote sensing images contain a wide variety of objects with large scale variations, and often contain multiple semantically similar or spatially adjacent objects. Traditional segmentation methods based on the single-target assumption struggle to support "multi-target" or "regional group" representations, which can easily lead to missed matches or positioning errors. Second, most existing methods rely on structured, template-based text representations for image-text alignment, making them difficult to adapt to the complex expressions of position, attributes, and quantity in natural language, limiting flexible human-computer interaction capabilities. Finally, in the actual implementation of remote sensing tasks, due to complex image backgrounds, fuzzy object boundaries, and strong interference, existing models lack generalization and robustness in open remote sensing scenes, prone to misidentification and false responses, affecting segmentation accuracy and system reliability. Summary of the Invention

[0005] In order to solve the technical problems existing in the above-mentioned prior art, the present invention provides a method and system for remote sensing image reference segmentation guided by free text, and the technical solution is as follows:

[0006] In one aspect, a free text-guided remote sensing image referent segmentation method is provided, the method comprising:

[0007] S1. Collect and preprocess remote sensing image data;

[0008] S2. Perform instance-level object mask annotation and other processing on the preprocessed image data to obtain various labels, construct a diverse natural language description of the article, and construct a data sample including images, text, and various labels;

[0009] S3, inputting the data sample into and training a region-relationship driven image-text segmentation model, wherein the region-relationship driven image-text segmentation model includes a dynamic association visual encoder, a pixel-level decoder, a context-related text encoder, a region-relationship modeling module, and a goal-oriented joint decoder module;

[0010] The dynamic association visual encoder performs multi-scale perception and dynamic response enhancement on the input image data to generate multi-scale visual features with spatial structure information ;

[0011] The pixel-level decoder decodes the multi-scale visual features Perform pixel-level decoding and output image mask information including instances of each category of the image ;

[0012] The context-related text encoder performs semantic modeling on the input text, comprehensively extracts various key information contained therein, and generates attribute-object information with context structure perception capability. ;

[0013] The regional relationship modeling module and Perform region-visual modeling interaction and region-language modeling interaction respectively, gradually integrate various semantic information, improve the model's understanding and modeling capabilities of complex expressions, and obtain regional filters and regional association characteristics ;

[0014] The target-oriented joint decoder 、 and Perform joint decoding to determine whether the input text has a true semantic match with the image, and perform multi-target merging and non-target diagnosis to achieve the multi-head prediction output of the model: target mask , regional probability and target presence detection ;

[0015] S4. Use the trained regional relationship to drive the image and text segmentation model to segment the remote sensing image.

[0016] Optionally, the S2 specifically includes:

[0017] Perform instance-level object mask annotation on the preprocessed image to obtain pixel mask labels ;

[0018] The pixel mask is divided into spatial regions, and the pixel-level annotations in each region are compressed into a value through average pooling to represent the probability that each region in the image includes a specific target category, which is used as the region probability label , to reflect the spatial position of the target in the image and its distribution trend, which is used as a guiding signal in the subsequent training process to help the model focus on the key areas;

[0019] Count whether each image contains a certain type of target. If the pixel label of the corresponding category in the image is not empty, mark this category as "1"; otherwise, mark it as "0" to obtain the target existence label. , used to indicate the presence of each target category in the image, serving as an auxiliary supervisory signal in the subsequent training process;

[0020] Constructing diverse natural language descriptions of this article. Text design covers six types of expression structures, including counting and ordinal expressions, nested logical structure expressions, multi-target attribute differentiation expressions, complex relationship interaction expressions, irrelevant object definitions, and deceptive attribute design. This fully covers challenging description forms in remote sensing scenarios and enhances the model's understanding and discrimination capabilities of natural language.

[0021] Build get including image ,text , pixel mask label , regional probability label and target presence tag data sample.

[0022] Optionally, the dynamic association visual encoder first performs preliminary projection and normalization on the image features through latent space affine operation to obtain the original image encoding , improve the expression stability of input;

[0023] Then, multiple convolution kernels with different receptive fields are used to perform multi-way parallel convolution operations on image features to capture structural information and spatial context at different scales. The multi-scale features output by the multi-way parallel convolution are averaged point by point along the channel direction at each spatial position and fused into a unified feature, which not only preserves scale diversity but also avoids interference caused by channel redundancy.

[0024] The fused unified features are fed into the Sigmoid activation function to generate pixel-level response weights to represent the significance of pixels in each region in the global field of view. The pixel-level response weights are multiplied point by point with the multi-scale features to achieve feature enhancement and suppression, and obtain semantic enhancement features. ;

[0025] Then the Input visual gating unit, the visual gating unit combines the With the learnable transformation matrix, nonlinear activation is applied along the channel dimension, and corresponding adjustment weights are generated for each channel to constrain the , and finally output multi-scale visual features , the formula is as follows:

[0026]

[0027] in represents the Sigmoid activation function, represents the learnable transformation matrix, Indicates element-wise multiplication of each channel.

[0028] Optionally, the context-sensitive text encoder first maps the input text into a vector representation through semantic embedding, and introduces TextBlob to perform part-of-speech analysis to extract key grammatical structures;

[0029] Then the text vector and key grammatical structure are respectively passed through the text global feature extractor and text local feature extractor built based on the BERT model to model long-distance dependencies and local semantic details, comprehensively capture the target information and spatial relationships in the description, and obtain local features. and global features ;

[0030] described and Through the gated weighted fusion module, dynamic fusion of multi-granularity semantic features is achieved, including:

[0031] The and Splicing, through a learnable linear transformation matrix Extract the fusion weight and normalize it through the Sigmoid activation function to generate the weight gating factor, which is used to control the fusion ratio of the two features and finally output the fused attribute-object information , to achieve the adaptive combination of multi-granularity information at the semantic level. The formula is as follows:

[0032]

[0033] in represents the Sigmoid activation function, represents the learnable transformation matrix.

[0034] Optionally, the regional relationship modeling module includes a regional-visual cross-fusion RVI submodule and a regional-language cross-fusion RLI submodule;

[0035] The RVI submodule dynamically extracts features of semantically relevant regions in an image through a region-level attention mechanism, including:

[0036] The original input image is fixed in size. Cut and get Representative regions, randomly initialize the learnable region query vector for the segmented representative regions and with the Combined with the learnable weight configuration, the region-image association matrix is ​​calculated , the formula is as follows:

[0037]

[0038] in is the learnable parameter weight;

[0039] Based on the , through the linear perception layer to perform nonlinear transformation and information compression, the regional filter is obtained , It integrates the features of the region at different scales and positions, as well as important information related to it in the global context, reflecting the response strength and content matching of each region in the image semantic structure, helping to distinguish regions related to the target description from irrelevant background for subsequent mask filtering;

[0040] According to the , extracting regional visual perception features from images , the formula is as follows:

[0041]

[0042] in is the linear transformation weight;

[0043] The RLI submodule models the dependencies between regions through the autocorrelation fitting mechanism to obtain regional relationship perception features with global correlation modeling. , the relationship between region and language is modeled through the cross-correlation fitting mechanism, and the regional language perception features that combine attribute-object information are obtained , then, the 、 as well as Weighted fusion is performed to obtain regional association features that have both spatial structure, context association and language guidance capabilities ;

[0044] The autocorrelation fitting mechanism is described as follows: The system is split by region, and after linear transformation, self-attention calculation is performed between each region to generate the similarity matrix between regions. The similarity matrix is ​​normalized by Softmax operation, and then the inter-region attention weight matrix is ​​obtained by feedforward neural network. The inter-region attention weight matrix and Point product fusion, we get the ;

[0045] The cross-correlation fitting mechanism, the and After their respective linear transformations, the region-to-language response matrix is ​​generated by cross-attention calculation. The response matrix is ​​normalized and then captured by a feedforward neural network to capture nonlinear features. The region-language attention matrix is ​​obtained, which represents the strength of the association between the region and each word in the language description. The region-language attention matrix is ​​combined with Point product fusion, we get the .

[0046] Optionally, the target-oriented joint decoder jointly decodes the image instance pixel mask information and the regional semantic features, effectively integrating regional semantics, location awareness, and language guidance information, including:

[0047] Receive instance-oriented image mask information Two types of output generated by regional relationship modeling: regional association features and regional filters As input;

[0048] For the nth region, its corresponding regional association feature The confidence that this area includes the target is output through the linear perception layer, and the regional filter With image mask information Perform point-by-point multiplication to generate a regional filtering mask that is more consistent with the foreground response , indicating its specific location and range in the image, and the regional filter mask of all regions Perform weighted aggregation based on their respective confidence levels and predict the overall target mask output ;

[0049] Based on the confidence level of each region including the target, the region probability is output as a representation of the scope of the target's existence;

[0050] Through Perform global average pooling to reduce the dimension to a fixed dimension, and use the linear perception layer to obtain the model's target presence judgment on whether the image contains a specified object .

[0051] Optionally, the loss function of the region relationship driven image-text segmentation model includes:

[0052] Predicted object mask The pixel-level segmentation loss calculated by comparing with the pixel label is used to supervise the consistency between the pixel-level target mask and the true label. The standard two-class cross entropy loss is used, which is denoted as :

[0053]

[0054] in The model predicts the The probability of the target existing in pixels, Indicates the The true mask label of pixels;

[0055] Predicted regional probability Region-level cross entropy loss calculated by comparing with region labels , to evaluate whether the model can correctly identify which areas contain targets:

[0056]

[0057] in Indicates the probability that the model predicts whether the nth region contains the target, Represents the true probability label of the nth region;

[0058] Predicted target existence judgment The target discrimination classification loss calculated by comparing with the target existence label is binary classification loss, which is recorded as :

[0059]

[0060] in Indicates the probability that the model predicts whether the entire image contains the target, Indicates the presence of an object label in the image;

[0061] The three loss items mentioned above are: 、 、 , together constitute the final loss function, which is used to simultaneously optimize the image region segmentation ability, regional target discrimination ability and image-text semantic consistency modeling ability.

[0062] In another aspect, a free text-guided remote sensing image referent segmentation system is provided, the system comprising:

[0063] Collection and preprocessing module, used to collect and preprocess remote sensing image data;

[0064] A construction module is used to perform instance-level object mask annotation and other processing on the pre-processed image data to obtain various labels, construct a diverse natural language description of the article, and construct data samples including images, text, and various labels;

[0065] A training module, configured to input the data samples and train a region-relationship-driven image-text segmentation model, wherein the region-relationship-driven image-text segmentation model includes a dynamic association visual encoder, a pixel-level decoder, a context-related text encoder, a region-relationship modeling module, and a goal-oriented joint decoder module;

[0066] The dynamic association visual encoder performs multi-scale perception and dynamic response enhancement on the input image data to generate multi-scale visual features with spatial structure information ;

[0067] The pixel-level decoder decodes the multi-scale visual features Perform pixel-level decoding and output image mask information including instances of each category of the image ;

[0068] The context-related text encoder performs semantic modeling on the input text, comprehensively extracts various key information contained therein, and generates attribute-object information with context structure perception capability. ;

[0069] The regional relationship modeling module and Perform region-visual modeling interaction and region-language modeling interaction respectively, gradually integrate various semantic information, improve the model's understanding and modeling capabilities of complex expressions, and obtain regional filters and regional association characteristics ;

[0070] The target-oriented joint decoder 、 and Perform joint decoding to determine whether the input text has a true semantic match with the image, and perform multi-target merging and non-target diagnosis to achieve the multi-head prediction output of the model: target mask , regional probability and target presence detection ;

[0071] The segmentation module is used to drive the image and text segmentation model using the trained regional relationships to segment the remote sensing images.

[0072] On the other hand, an electronic device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the above-mentioned free text-guided remote sensing image reference segmentation method.

[0073] On the other hand, a computer-readable storage medium is provided, wherein the storage medium stores at least one instruction, and the at least one instruction is loaded and executed by a processor to implement the above-mentioned free text-guided remote sensing image reference segmentation method.

[0074] The beneficial effects brought about by the technical solution provided by the present invention include at least:

[0075] 1) The regional relationship modeling mechanism designed in this paper realizes contextual interaction modeling among multiple targets by constructing an explicit semantic association graph of regions in the image. It supports the overall parsing and positioning of general and set descriptions in natural language, and significantly improves the joint recognition and segmentation accuracy of multiple instance targets in complex remote sensing scenarios.

[0076] 2) The dynamic associative visual encoder and context-associated text encoder designed in this paper can flexibly adjust the regional perception range according to the input semantics and integrate multi-scale contextual semantic features to achieve accurate modeling and positioning of targets of different scales, significantly enhancing the accuracy, stability and generalization ability in open remote sensing scenarios.

[0077] 3) The goal-oriented joint decoder designed in this paper introduces a region-level cross-modal matching and existence judgment mechanism, which can effectively identify whether the input text has actual visual correspondence, avoid erroneous responses to irrelevant descriptions, and significantly improve the discrimination ability and stability under free natural language expression.

[0078] 4) This paper adopts a multi-task joint training strategy to perform weighted optimization on the losses of the three tasks of target mask, regional target probability, and target existence discrimination, ensuring that the model achieves high precision in multiple tasks. This optimization strategy can improve the accuracy of target positioning, segmentation, and discrimination while maintaining system stability, enabling the model to perform well in various complex remote sensing image analysis tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0080] Figure 1 This is a flow chart of a free text guided remote sensing image reference segmentation method provided by an embodiment of the present invention;

[0081] Figure 2 This is an overall block diagram of a free text-guided remote sensing image reference segmentation method provided by an embodiment of the present invention;

[0082] Figure 3 This is a schematic diagram of constructing a data sample provided by an embodiment of the present invention;

[0083] Figure 4 This is a structural block diagram of a region relationship driven image and text segmentation model provided by an embodiment of the present invention;

[0084] Figure 5 This is a structural block diagram of a dynamic association visual encoder provided by an embodiment of the present invention;

[0085] Figure 6 This is a structural block diagram of a context-sensitive text encoder provided by an embodiment of the present invention;

[0086] Figure 7 This is a structural block diagram of a regional relationship modeling module provided by an embodiment of the present invention;

[0087] Figure 8 Schematic diagram of the autocorrelation fitting mechanism provided by an embodiment of the present invention;

[0088] Figure 9 is a schematic diagram of a cross-correlation fitting mechanism provided by an embodiment of the present invention;

[0089] Figure 10 is a structural block diagram of a goal-oriented joint decoder provided by an embodiment of the present invention;

[0090] Figure 11 This is a block diagram of a remote sensing image reference segmentation system guided by free text provided by an embodiment of the present invention;

[0091] Figure 12 It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0092] In order to make the technical problems, technical solutions and advantages to be solved by the present invention clearer, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.

[0093] An embodiment of the present invention provides a free text-guided remote sensing image referent segmentation method, which can be implemented by an electronic device, which can be a terminal or a server. Figure 1 The flow chart of this method is shown in FIG. Figure 2 The figure shows the overall block diagram of this method, which aims to solve the problems of multi-target expression being difficult to aggregate, misjudgment of non-target descriptions, and semantic segmentation of multi-scale objects in remote sensing image segmentation. The specific processing flow can include the following steps:

[0094] S1. Collect and preprocess remote sensing image data;

[0095] In the embodiment of the present invention, a drone equipped with a high-resolution visible light sensor is used to conduct low-altitude aerial photography of typical areas such as urban buildings, rural roads, farmland, and water bodies, and collect remote sensing image data with clear structure and rich semantics. The images are original color images with a uniform size of 4000×3000 pixels and a resolution better than 0.2 meters / pixel. They can accurately present the outlines and positional relationships of common landforms. During the collection process, the flight altitude and angle are controlled to ensure that the images are unobstructed and have minimal distortion, meeting the needs of subsequent annotation and text construction.

[0096] After completing the remote sensing image data acquisition, the embodiment of the present invention pre-processes the original image, including unifying the format specification and size cropping, and standardizing it to 512×512 pixels to meet the model input requirements.

[0097] S2. Perform instance-level object mask annotation and other processing on the preprocessed image data to obtain various labels, construct a diverse natural language description of the article, and construct a data sample including images, text, and various labels;

[0098] Alternatively, as Figure 3 As shown, the S2 specifically includes:

[0099] Perform instance-level object mask annotation on the preprocessed image (accurately mark the categories and boundaries of objects such as buildings, roads, water bodies, vegetation, vehicles, etc. to ensure clear semantics and complete structure) to obtain pixel mask labels ;

[0100] The pixel mask is divided into spatial regions, and the pixel-level annotations in each region are compressed into a value through average pooling to represent the probability that each region in the image includes a specific target category, which is used as the region probability label , to reflect the spatial position of the target in the image and its distribution trend, which is used as a guiding signal in the subsequent training process to help the model focus on the key areas;

[0101] Count whether each image contains a certain type of target. If the pixel label of the corresponding category in the image is not empty, mark this category as "1"; otherwise, mark it as "0" to obtain the target existence label. , used to indicate the presence of each target category in the image, serving as an auxiliary supervisory signal in the subsequent training process;

[0102] Construct diverse natural language descriptions of this article (free text). The text design covers six types of expression structures, including counting and ordinal expressions (used for target quantity and sorting and positioning), nested logical structure expressions (reflecting parallel and exclusive relationships between targets), multi-target attribute differentiation expressions (describing targets with multiple attribute similarities and differences), complex relationship interaction expressions (involving spatial or semantic attachment between targets), irrelevant object definitions (constructing expressions that are partially related but have no corresponding targets), and deceptive attribute design (introducing interfering descriptions to enhance discriminative robustness). This fully covers challenging description forms in remote sensing scenarios and enhances the model's understanding and discrimination capabilities of natural language.

[0103] Build get including image ,text , pixel mask label , regional probability label and target presence tag data sample.

[0104] In addition, after completing the data sample construction, the embodiment of the present invention divides the data into a training set and a validation set in a ratio of 4:1, and reasonably organizes different types of text-image pairs, including three types of samples: one-to-one, one-to-many, and one-to-zero. For example, in the one-to-one type sample, the text description is "the red-roofed building in the lower right corner", which only corresponds to a single ground object target in the image; in the one-to-many type sample, the description is "white cars on both sides of the road", involving multiple semantically related instances in the image; in the one-to-zero type sample, the description is such as "the pedestrian in blue clothes in the picture", but there is no corresponding target in the image. During the division process, the distribution of various types of samples in the training set and the validation set is kept balanced to enhance the generalization ability and evaluation stability of the model. In addition, a background random masking strategy is introduced. During the training stage, non-target areas in the image are randomly selected, and block occlusion is performed in a fixed-size sliding window manner, and replaced with mean filling or noise perturbation, thereby suppressing background interference and enhancing the model's semantic focus and discrimination ability on key target areas.

[0105] S3, input the data sample and train the regional relationship driven image and text segmentation model, the regional relationship driven image and text segmentation model includes a dynamic association visual encoder, a pixel level decoder, a context association text encoder, a regional relationship modeling module and a target-oriented joint decoder module, such as Figure 4 As shown;

[0106] The dynamic association visual encoder performs multi-scale perception and dynamic response enhancement on the input image data to generate multi-scale visual features with spatial structure information ;

[0107] The pixel-level decoder decodes the multi-scale visual features Perform pixel-level decoding and output image mask information including instances of each category of the image ;

[0108] The context-related text encoder performs semantic modeling on the input text, comprehensively extracts various key information contained therein, and generates attribute-object information with context structure perception capability. ;

[0109] The regional relationship modeling module and Perform region-visual modeling interaction and region-language modeling interaction respectively, gradually integrate various semantic information, improve the model's understanding and modeling capabilities of complex expressions, and obtain regional filters and regional association characteristics ;

[0110] The target-oriented joint decoder 、 and Perform joint decoding to determine whether the input text has a true semantic match with the image, and perform multi-target merging and non-target diagnosis to achieve the multi-head prediction output of the model: target mask , regional probability and target presence detection ;

[0111] Alternatively, as Figure 5 As shown in the figure, the dynamic association visual encoder first performs preliminary projection and normalization on the image features through latent space affine operation to obtain the original image encoding , improve the expression stability of input;

[0112] Then, multiple convolution kernels with different receptive fields (including 3×3, 5×5, and 7×7) are used to perform multi-way parallel convolution operations on image features to capture structural information and spatial context at different scales. The multi-scale features output by the multi-way parallel convolution are averaged point by point along the channel direction at each spatial position and fused into a unified feature, preserving scale diversity while avoiding interference caused by channel redundancy.

[0113] The fused unified features are fed into the Sigmoid activation function to generate pixel-level response weights to represent the significance of pixels in each region in the global field of view. The pixel-level response weights are multiplied point by point with the multi-scale features to achieve feature enhancement and suppression, and obtain semantic enhancement features. ;

[0114] Then the Input visual gating unit, the visual gating unit combines the With the learnable transformation matrix, nonlinear activation is applied along the channel dimension, and corresponding adjustment weights are generated for each channel to constrain the , and finally output multi-scale visual features , the formula is as follows:

[0115]

[0116] in represents the Sigmoid activation function, represents the learnable transformation matrix, Indicates element-wise multiplication of each channel.

[0117] The visual gating unit of the embodiment of the present invention can dynamically adjust the channel activation distribution of features according to the guidance information, strengthen the feature dimensions that are consistent with the semantics of the language description, thereby effectively improving the cross-modal alignment capability and discrimination performance.

[0118] Alternatively, as Figure 6 As shown in FIG, the context-sensitive text encoder first maps the input text into a vector representation through semantic embedding, and introduces TextBlob to perform part-of-speech analysis and extract key grammatical structures;

[0119] Then the text vector and key grammatical structure are respectively passed through the text global feature extractor and text local feature extractor built based on the BERT model to model long-distance dependencies and local semantic details, comprehensively capture the target information and spatial relationships in the description, and obtain local features. and global features ;

[0120] described and Through the gated weighted fusion module, dynamic fusion of multi-granularity semantic features is achieved, including:

[0121] The and Splicing, through a learnable linear transformation matrix Extract the fusion weight and normalize it through the Sigmoid activation function to generate the weight gating factor, which is used to control the fusion ratio of the two features and finally output the fused attribute-object information , to achieve the adaptive combination of multi-granularity information at the semantic level. The formula is as follows:

[0122]

[0123] in represents the Sigmoid activation function, represents the learnable transformation matrix.

[0124] Alternatively, as Figure 7 As shown, the regional relationship modeling module includes a region-vision integration (RVI) submodule and a region-language integration (RLI) submodule;

[0125] The RVI submodule dynamically extracts features of semantically relevant regions in an image through a region-level attention mechanism, including:

[0126] The original input image is fixed in size. Cut and get Representative regions, randomly initialize the learnable region query vector for the segmented representative regions and with the Combined with the learnable weight configuration, the region-image association matrix is ​​calculated , the formula is as follows:

[0127]

[0128] in is the learnable parameter weight;

[0129] Based on the , through the linear perception layer to perform nonlinear transformation and information compression, the regional filter is obtained , It integrates the features of the region at different scales and positions, as well as important information related to it in the global context, reflecting the response strength and content matching of each region in the image semantic structure, helping to distinguish regions related to the target description from irrelevant background for subsequent mask filtering;

[0130] According to the , extracting regional visual perception features from images , the formula is as follows:

[0131]

[0132] in is the linear transformation weight;

[0133] The above mechanism enables each regional feature to be dynamically aggregated from the entire map, which has higher flexibility and adaptability.

[0134] The RLI submodule models the dependencies between regions through the autocorrelation fitting mechanism to obtain regional relationship perception features with global correlation modeling. , the relationship between region and language is modeled through the cross-correlation fitting mechanism, and the regional language perception features that combine attribute-object information are obtained , then, the 、 as well as Weighted fusion is performed to obtain regional association features that have both spatial structure, context association and language guidance capabilities ;

[0135] The autocorrelation fitting mechanism, such as Figure 8 As shown, the The system is split by region, and after linear transformation, self-attention calculation is performed between each region to generate the similarity matrix between regions. The similarity matrix is ​​normalized by Softmax operation, and then the inter-region attention weight matrix is ​​obtained by feedforward neural network. The inter-region attention weight matrix and Point product fusion, we get the ;

[0136] This mechanism effectively captures the semantic associations between multiple targets and regions in complex backgrounds in remote sensing images, helping to overcome the problem of semantic ambiguity.

[0137] The cross-correlation fitting mechanism, such as Figure 9 As shown, the and After their respective linear transformations, the region-to-language response matrix is ​​generated by cross-attention calculation. The response matrix is ​​normalized and then captured by a feedforward neural network to capture nonlinear features. The region-language attention matrix is ​​obtained, which represents the strength of the association between the region and each word in the language description. The region-language attention matrix is ​​combined with Point product fusion, we get the .

[0138] This mechanism enhances the model's ability to semantically align language descriptions with image content, thereby improving the accuracy of target localization and segmentation.

[0139] Alternatively, as Figure 10As shown in Figure 2, the target-oriented joint decoder jointly decodes the image instance pixel mask information and the regional semantic features, effectively integrating regional semantics, location awareness, and language guidance information, including:

[0140] Receive instance-oriented image mask information Two types of output generated by regional relationship modeling: regional association features and regional filters As input;

[0141] For the nth region, its corresponding regional association feature The confidence that this area includes the target is output through the linear perception layer, and the regional filter With image mask information Perform point-by-point multiplication to generate a regional filtering mask that is more consistent with the foreground response , indicating its specific location and range in the image, and the regional filter mask of all regions Perform weighted aggregation based on their respective confidence levels and predict the overall target mask output ;

[0142] Based on the confidence level of each region including the target, the region probability is output as a representation of the scope of the target's existence;

[0143] Through Perform global average pooling to reduce the dimension to a fixed dimension, and use the linear perception layer to obtain the model's target presence judgment on whether the image contains a specified object .

[0144] This joint decoding mechanism effectively integrates regional semantics, location perception and language guidance information, has strong spatial reasoning and expression matching capabilities, and significantly improves the segmentation accuracy and robustness in multi-target and multi-relationship scenarios in remote sensing images.

[0145] Optionally, the loss function of the region relationship driven image-text segmentation model includes:

[0146] Predicted object mask The pixel-level segmentation loss calculated by comparing with the pixel label is used to supervise the consistency between the pixel-level target mask and the true label. The standard two-class cross entropy loss is used, which is denoted as :

[0147]

[0148] in The model predicts the The probability of the target existing in pixels, Indicates the The true mask label of pixels;

[0149] Predicted regional probability Region-level cross entropy loss calculated by comparing with region labels , to evaluate whether the model can correctly identify which areas contain targets:

[0150]

[0151] in Indicates the probability that the model predicts whether the nth region contains the target, Represents the true probability label of the nth region;

[0152] Predicted target existence judgment The target discrimination classification loss calculated by comparing with the target existence label is binary classification loss, which is recorded as :

[0153]

[0154] in Indicates the probability that the model predicts whether the entire image contains the target, Indicates the presence of an object label in the image;

[0155] The three loss items mentioned above are: 、 、 , together constitute the final loss function, which is used to simultaneously optimize the image region segmentation ability, regional target discrimination ability and image-text semantic consistency modeling ability.

[0156] S4. Use the trained regional relationship to drive the image and text segmentation model to segment the remote sensing image.

[0157] like Figure 11 As shown, an embodiment of the present invention further provides a remote sensing image reference segmentation system guided by free text, the system comprising:

[0158] A collection and preprocessing module 1110 is used to collect and preprocess remote sensing image data;

[0159] A construction module 1120 is configured to perform instance-level object mask annotation and other processing on the pre-processed image data to obtain various labels, construct a diverse natural language description of the text, and construct a data sample including images, text, and various labels;

[0160] A training module 1130 is configured to input the data samples and train a region-relationship driven image-text segmentation model, wherein the region-relationship driven image-text segmentation model includes a dynamic association visual encoder, a pixel-level decoder, a context-related text encoder, a region-relationship modeling module, and a goal-oriented joint decoder module;

[0161] The dynamic association visual encoder performs multi-scale perception and dynamic response enhancement on the input image data to generate multi-scale visual features with spatial structure information ;

[0162] The pixel-level decoder decodes the multi-scale visual features Perform pixel-level decoding and output image mask information including instances of each category of the image ;

[0163] The context-related text encoder performs semantic modeling on the input text, comprehensively extracts various key information contained therein, and generates attribute-object information with context structure perception capability. ;

[0164] The regional relationship modeling module and Perform region-visual modeling interaction and region-language modeling interaction respectively, gradually integrate various semantic information, improve the model's understanding and modeling capabilities of complex expressions, and obtain regional filters and regional association characteristics ;

[0165] The target-oriented joint decoder 、 and Perform joint decoding to determine whether the input text has a true semantic match with the image, and perform multi-target merging and non-target diagnosis to achieve the multi-head prediction output of the model: target mask , regional probability and target presence detection ;

[0166] The segmentation module 1140 is used to drive the image and text segmentation model using the trained regional relationship to segment the remote sensing image.

[0167] The free text-guided remote sensing image reference segmentation system provided in an embodiment of the present invention has a functional structure corresponding to the free text-guided remote sensing image reference segmentation method provided in an embodiment of the present invention, which will not be described in detail here.

[0168] Figure 12It is a structural diagram of an electronic device 1200 provided in an embodiment of the present invention. The electronic device 1200 may have relatively large differences due to different configurations or performances, and may include one or more processors (central processing units, CPU) 1201 and one or more memories 1202, wherein the memory 1202 stores at least one instruction, and the at least one instruction is loaded and executed by the processor 1201 to implement the steps of the above-mentioned free text-guided remote sensing image reference segmentation method.

[0169] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions. The instructions are executable by a processor in a terminal to implement the above-described free-text-guided remote sensing image referent segmentation method. For example, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, or optical data storage device.

[0170] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.

[0171] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A free text guided remote sensing image reference segmentation method, characterized in that: The method comprises: S1. Collect and preprocess remote sensing image data; S2. Perform instance-level object mask annotation and other processing on the preprocessed image data to obtain various labels, construct a diverse natural language description of the article, and construct a data sample including images, text, and various labels; S3, inputting the data sample into and training a region-relationship driven image-text segmentation model, wherein the region-relationship driven image-text segmentation model includes a dynamic association visual encoder, a pixel-level decoder, a context-related text encoder, a region-relationship modeling module, and a goal-oriented joint decoder module; The dynamic association visual encoder performs multi-scale perception and dynamic response enhancement on the input image data to generate multi-scale visual features with spatial structure information ; The pixel-level decoder decodes the multi-scale visual features Perform pixel-level decoding and output image mask information including instances of each category of the image ; The context-related text encoder performs semantic modeling on the input text, comprehensively extracts various key information contained therein, and generates attribute-object information with context structure perception capability. ; The regional relationship modeling module and Perform region-visual modeling interaction and region-language modeling interaction respectively, gradually integrate various semantic information, improve the model's understanding and modeling capabilities of complex expressions, and obtain regional filters and regional association characteristics ; The target-oriented joint decoder 、 and Perform joint decoding to determine whether the input text has a true semantic match with the image, and perform multi-target merging and non-target diagnosis to achieve the multi-head prediction output of the model: target mask , regional probability and target presence detection ; S4. Use the trained regional relationship to drive the image and text segmentation model to segment the remote sensing image.

2. The method according to claim 1, characterized in that Said S2 specifically includes: Perform instance-level object mask annotation on the preprocessed image to obtain pixel mask labels ; The pixel mask is divided into spatial regions, and the pixel-level annotations in each region are compressed into a value through average pooling to represent the probability that each region in the image includes a specific target category, which is used as the region probability label , to reflect the spatial position of the target in the image and its distribution trend, which is used as a guiding signal in the subsequent training process to help the model focus on the key areas; Count whether each image contains a certain type of target. If the pixel label of the corresponding category in the image is not empty, mark this category as "1"; otherwise, mark it as "0" to obtain the target existence label. , used to indicate the presence of each target category in the image, serving as an auxiliary supervisory signal in the subsequent training process; Constructing diverse natural language descriptions of this article. Text design covers six types of expression structures, including counting and ordinal expressions, nested logical structure expressions, multi-target attribute differentiation expressions, complex relationship interaction expressions, irrelevant object definitions, and deceptive attribute design. This fully covers challenging description forms in remote sensing scenarios and enhances the model's understanding and discrimination capabilities of natural language. Build get including image ,text , pixel mask label , regional probability label and target presence tag data sample.

3. The method according to claim 1, characterized in that The dynamic association visual encoder first performs preliminary projection and normalization on the image features through latent space affine operation to obtain the original image encoding , improve the expression stability of input; Then, multiple convolution kernels with different receptive fields are used to perform multi-way parallel convolution operations on image features to capture structural information and spatial context at different scales. The multi-scale features output by the multi-way parallel convolution are averaged point by point along the channel direction at each spatial position and fused into a unified feature, which not only preserves scale diversity but also avoids interference caused by channel redundancy. The fused unified features are fed into the Sigmoid activation function to generate pixel-level response weights to represent the significance of pixels in each region in the global field of view. The pixel-level response weights are multiplied point by point with the multi-scale features to achieve feature enhancement and suppression, and obtain semantic enhancement features. ; Then the Input visual gating unit, the visual gating unit combines the With the learnable transformation matrix, nonlinear activation is applied along the channel dimension, and corresponding adjustment weights are generated for each channel to constrain the , and finally output multi-scale visual features , the formula is as follows: ; in represents the Sigmoid activation function, represents the learnable transformation matrix, Indicates element-wise multiplication of each channel.

4. The method according to claim 1, wherein The context-sensitive text encoder first maps the input text into a vector representation through semantic embedding, and introduces TextBlob to perform part-of-speech analysis and extract key grammatical structures; Then the text vector and key grammatical structure are respectively passed through the text global feature extractor and text local feature extractor built based on the BERT model to model long-distance dependencies and local semantic details, comprehensively capture the target information and spatial relationships in the description, and obtain local features. and global features ; described and Through the gated weighted fusion module, dynamic fusion of multi-granularity semantic features is achieved, including: The and Splicing, through a learnable linear transformation matrix Extract the fusion weight and normalize it through the Sigmoid activation function to generate the weight gating factor, which is used to control the fusion ratio of the two features and finally output the fused attribute-object information , to achieve the adaptive combination of multi-granularity information at the semantic level. The formula is as follows: ; in represents the Sigmoid activation function, represents the learnable transformation matrix.

5. The method according to claim 1, characterized in that The regional relationship modeling module includes a regional-visual cross fusion RVI submodule and a regional-language cross fusion RLI submodule; The RVI submodule dynamically extracts features of semantically relevant regions in an image through a region-level attention mechanism, including: The original input image is fixed in size. Cut and get Representative regions are randomly initialized for the segmented representative regions to learn the region query vector and with the Combined with the learnable weight configuration, the region-image association matrix is ​​calculated , the formula is as follows: ; in is the learnable parameter weight; Based on the , through the linear perception layer to perform nonlinear transformation and information compression, the regional filter is obtained , It integrates the features of the region at different scales and positions, as well as important information related to it in the global context, reflecting the response strength and content matching of each region in the image semantic structure, helping to distinguish regions related to the target description from irrelevant background for subsequent mask filtering; According to the , extracting regional visual perception features from images , the formula is as follows: ; in is the linear transformation weight; The RLI submodule models the dependencies between regions through the autocorrelation fitting mechanism to obtain regional relationship perception features with global correlation modeling. , the relationship between region and language is modeled through the cross-correlation fitting mechanism, and the regional language perception features that combine attribute-object information are obtained , then, the 、 as well as Weighted fusion is performed to obtain regional association features that have both spatial structure, context association and language guidance capabilities ; The autocorrelation fitting mechanism is described as follows: The system is split by region, and after linear transformation, self-attention calculation is performed between each region to generate the similarity matrix between regions. The similarity matrix is ​​normalized by Softmax operation, and then the inter-region attention weight matrix is ​​obtained by feedforward neural network. The inter-region attention weight matrix and Point product fusion, we get the ; The cross-correlation fitting mechanism, the and After their respective linear transformations, the region-to-language response matrix is ​​generated by cross-attention calculation. The response matrix is ​​normalized and then captured by a feedforward neural network to capture nonlinear features. The region-language attention matrix is ​​obtained, which represents the strength of the association between the region and each word in the language description. The region-language attention matrix is ​​combined with Point product fusion, we get the .

6. The method according to claim 1, characterized in that The target-oriented joint decoder jointly decodes the image instance pixel mask information and regional semantic features, effectively integrating regional semantics, location awareness, and language guidance information, including: Receive instance-oriented image mask information Two types of output generated by regional relationship modeling: regional association features and regional filters As input; For the nth region, its corresponding regional association feature The confidence that this area includes the target is output through the linear perception layer, and the regional filter With image mask information Perform point-by-point multiplication to generate a regional filtering mask that is more consistent with the foreground response , indicating its specific location and range in the image, and the regional filter mask of all regions Perform weighted aggregation based on their respective confidence levels and predict the overall target mask output ; Based on the confidence level of each region including the target, the region probability is output as a representation of the scope of the target's existence; Through Perform global average pooling to reduce the dimension to a fixed dimension, and use the linear perception layer to obtain the model's target presence judgment on whether the image contains a specified object .

7. The method according to claim 1, characterized in that The loss function of the region relationship driven image and text segmentation model includes: Predicted object mask The pixel-level segmentation loss calculated by comparing with the pixel label is used to supervise the consistency between the pixel-level target mask and the true label. The standard two-class cross entropy loss is used, which is denoted as : ; in The model predicts the The probability of the target existing in pixels, Indicates the The true mask label of pixels; Predicted regional probability Region-level cross entropy loss calculated by comparing with region labels , to evaluate whether the model can correctly identify which areas contain targets: ; in Indicates the probability that the model predicts whether the nth region contains the target, Represents the true probability label of the nth region; Predicted target existence judgment The target discrimination classification loss calculated by comparing with the target existence label is binary classification loss, which is recorded as : ; in Indicates the probability that the model predicts whether the entire image contains the target, Indicates the presence of an object label in the image; The three loss items mentioned above are: 、 、 , together constitute the final loss function, which is used to simultaneously optimize the image region segmentation ability, regional target discrimination ability and image-text semantic consistency modeling ability.

8. A free text guided remote sensing image reference segmentation system, characterized by: The system comprises: Collection and preprocessing module, used to collect and preprocess remote sensing image data; A construction module is used to perform instance-level object mask annotation and other processing on the pre-processed image data to obtain various labels, construct a diverse natural language description of the article, and construct data samples including images, text, and various labels; A training module, configured to input the data samples and train a region-relationship-driven image-text segmentation model, wherein the region-relationship-driven image-text segmentation model includes a dynamic association visual encoder, a pixel-level decoder, a context-related text encoder, a region-relationship modeling module, and a goal-oriented joint decoder module; The dynamic association visual encoder performs multi-scale perception and dynamic response enhancement on the input image data to generate multi-scale visual features with spatial structure information ; The pixel-level decoder decodes the multi-scale visual features Perform pixel-level decoding and output image mask information including instances of each category of the image ; The context-related text encoder performs semantic modeling on the input text, comprehensively extracts various key information contained therein, and generates attribute-object information with context structure perception capability. ; The regional relationship modeling module and Perform region-visual modeling interaction and region-language modeling interaction respectively, gradually integrate various semantic information, improve the model's understanding and modeling capabilities of complex expressions, and obtain regional filters and regional association characteristics ; The target-oriented joint decoder 、 and Perform joint decoding to determine whether the input text has a true semantic match with the image, and perform multi-target merging and non-target diagnosis to achieve the multi-head prediction output of the model: target mask , regional probability and target presence detection ; The segmentation module is used to drive the image and text segmentation model using the trained regional relationships to segment the remote sensing images.

9. An electronic device comprising a processor and a memory, wherein the memory stores at least one instruction, characterized in that: The at least one instruction is loaded and executed by the processor to implement the free text-guided remote sensing image reference segmentation method according to any one of claims 1 to 7.

10. A computer-readable storage medium, wherein at least one instruction is stored in the storage medium, characterized in that: The at least one instruction is loaded and executed by the processor to implement the free text guided remote sensing image reference segmentation method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Visual localization and anaphora segmentation method, system and device based on mask anaphora modeling and storage medium

    CN118734091A

  • Medical visual question and answer method and system based on multi-task modeling

    CN119202334A