A cross-modal remote sensing image-text retrieval method based on spatial layout perception
Patent Information
- Application Number
- CN202610681334.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-18
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-05-18
AI Technical Summary
现有方法由于缺乏对空间布局的显式建模,难以区分“建筑物包围水体”与“水体包围建筑物”等具有相反空间布局的图像,检索精度受限
[0076]有益效果:本发明的一种基于空间布局感知的跨模态遥感图文检索方法,通过构建图像语义图和文本语义图,引入空间布局约束,实现了图像区域与文本短语的细粒度匹配,有效解决了现有技术忽略物体空间关系导致的检索精度不足问题,显著提升了涉及空间布局描述的图文检索准确性。具体而言,本发明在图像侧采用目标检测网络提取多个地物目标区域,并为每个区域生成融合视觉内容与空间位置编码的节点特征,同时基于区域间的相对位置、距离、重叠程度及尺度差异构建边特征,形成结构化的图像语义图;在文本侧通过句法分析与关系抽取解析出实体短语及其空间关系谓词,并利用图神经网络将实体节点与关系边编码为文本语义图。显著提升了对复杂空间关系描述的检索鲁棒性。
Smart Images

Figure CN122200378B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of remote sensing image processing technology, and in particular to a cross-modal remote sensing image retrieval method based on spatial layout awareness. Background Technology
[0002] Remote sensing imagery, as an important geospatial data source, has wide applications in land resource surveys, urban planning, disaster monitoring, and military reconnaissance. With the rapid development of remote sensing satellite and UAV technologies, remote sensing image data is experiencing explosive growth. How to quickly and accurately retrieve the target images needed by users from massive amounts of remote sensing images has become a research hotspot in the field of remote sensing information processing. Cross-modal remote sensing image retrieval technology, by allowing users to query remote sensing images using natural language descriptions, greatly lowers the retrieval threshold and has significant theoretical implications and application prospects.
[0003] Existing cross-modal image-text retrieval methods are mainly based on deep learning frameworks, with CLIP (Contrastive Language-Image Pre-training) and its variants being typical examples. These methods usually employ a dual-tower architecture, utilizing a visual encoder and a text encoder respectively to extract global features of the image and global semantic features of the text. Through contrastive learning, they map the image and text features to a common semantic space, and finally achieve retrieval by calculating the cosine similarity of the feature vectors. This method has achieved good retrieval results in the natural image domain.
[0004] However, directly applying the above methods to remote sensing image retrieval has the following technical drawbacks:
[0005] First, remote sensing images are characterized by large scale, numerous targets, and complex backgrounds. These images often contain multiple ground features (such as buildings, roads, water bodies, and vegetation), and these features exhibit rich spatial relationships (e.g., "buildings are located on the left side of the road," "water bodies surround vegetation," etc.). Existing methods compress the entire image into a single feature vector using global pooling, severely losing fine-grained information and spatial layout structure of the targets, making it difficult to meet users' retrieval needs for descriptions of complex spatial relationships.
[0006] Second, the text descriptions of remote sensing images typically contain information about multiple entities and their relative positions, such as "the left side of the image is a residential area, the right side is farmland, and a river runs through the middle." Existing methods use global encoding for the text, which cannot effectively capture the spatial relationship semantics between multiple entities. As a result, when matching images and text, only the existence of the target can be guaranteed, but layout consistency cannot be guaranteed.
[0007] Third, remote sensing images often exhibit high visual similarity of similar ground features, and the spatial topological relationships between different ground features are crucial for distinguishing scene categories. Existing methods, lacking explicit modeling of spatial layouts, struggle to differentiate between images with opposite spatial layouts, such as "buildings surrounding water" and "water surrounding buildings," thus limiting retrieval accuracy.
[0008] In summary, there is an urgent need for a cross-modal remote sensing image and text retrieval method that can perceive spatial layout. By finely modeling the spatial relationships of ground objects in images, it can achieve accurate matching of images and text in both content and layout dimensions, thereby solving the problem of insufficient retrieval accuracy of existing technologies in remote sensing scenarios. Summary of the Invention
[0009] The purpose of this invention is to provide a cross-modal remote sensing image and text retrieval method based on spatial layout awareness. By finely modeling the spatial relationships of ground objects in images, it achieves accurate matching of images and text in both content and layout dimensions, significantly improving retrieval performance in complex remote sensing scenarios.
[0010] To achieve the above objectives, the present invention provides a cross-modal remote sensing image and text retrieval method based on spatial layout awareness, comprising:
[0011] S1. Acquire the image, extract multiple target regions from the image, and generate region features and spatial location codes for each target region;
[0012] S2. Obtain the text query, parse the multiple entity phrases in the text query and the spatial relationships between the entity phrases, and construct a text semantic graph;
[0013] S3. Construct an image semantic graph by using the multiple target regions as nodes and the spatial relationships between regions as edges;
[0014] S4. Calculate the node matching score and graph structure matching score of the image semantic graph and the text semantic graph through a cross-modal graph matching network, and fuse them to obtain the image-text similarity score;
[0015] S5. Output the search results based on the image-text similarity score.
[0016] Furthermore, the specific steps in S1 of acquiring the image, extracting multiple target regions from the image, and generating region features and spatial location codes for each target region include:
[0017] S1.1 Input Image A target detection network is used to extract K candidate target regions. Each region i outputs a visual feature vector. Bounding box coordinates After normalization, the processing formula is as follows:
[0018] ;
[0019] Where W and H are the original image width and height; , These are the x-coordinates of the left and right boundaries of the bounding box of the i-th target region in the original image, respectively. , Let y be the upper and lower boundaries of the bounding box of the i-th target region in the original image;
[0020] S1.2 To incorporate spatial information into the region representation, the normalized bounding box... Encoding is performed. Absolute position encoding is used to map the boundary coordinates into learnable embedding vectors, calculated as follows:
[0021] ;
[0022] MLP stands for Multilayer Perceptron. The dimension for location encoding;
[0023] S1.3. Fuse visual features with location codes to obtain the final region fusion features. The calculation formula is as follows:
[0024] ;
[0025] [ ; ] indicates a splicing operation.
[0026] Furthermore, the specific steps in S2 to obtain the text query, parse multiple entity phrases in the text query and the spatial relationships between the entity phrases, and construct a text semantic graph include:
[0027] S2.1 Using a pre-trained named entity recognition or noun block recognition model, let the extracted entity set be... Where P is the number of entities, for each entity phrase The semantic feature vector is extracted using a pre-trained language model (BERT), and the calculation formula is as follows:
[0028] ;
[0029] in, For the encoder of the language model, it can take CLS tags or average pooling output. For feature dimensions;
[0030] S2.2, Let the extracted spatial relation set be... , representing entities and There exists a spatial relationship between them. Each relationship can be represented as a triple. ,in For relational predicates, the relational predicates are mapped to learnable embedding vectors, and the calculation formula is as follows:
[0031] ;
[0032] in, For relational embedding functions;
[0033] S2.3, Node is Each node Let be the entity feature vector, and be the edge vector. To integrate node features into neighbor relationships and the global context, Graph Attention Networks (GAT) or Graph Convolutional Networks (GCN) can be used for message passing. The calculation formula is as follows:
[0034] ;
[0035] in, Let i be the set of neighboring nodes. Attention coefficient It is an aggregate function.
[0036] Furthermore, the specific steps in S3 for constructing an image semantic graph by using the multiple target regions as nodes and the spatial relationships between regions as edges include:
[0037] S3.1 Using the normalized bounding box Calculate the following components:
[0038] Center coordinates: ;
[0039] Relative center offset: ;
[0040] Normalized central distance: ;
[0041] Intersection over Union (IoU): ;
[0042] Relative area ratio: ;
[0043] Direction angle: ;
[0044] The above components are concatenated to form the original geometric feature vector:
[0045] ;
[0046] in, Let x and y be the center coordinates of region i; Let x and y be the center coordinates of region j; The horizontal and vertical offsets are relative to the center. Let i be the bounding box of regions i and j; , The area of the overlapping region between the two frames, and the total area covered by the two frames; Let i be the area of regions i and j; For the direction sine and cosine;
[0047] S3.2. The original geometric features are mapped to high-dimensional edge features using a multilayer perceptron (MLP). The calculation formula is as follows:
[0048] ;
[0049] in, Set it to 128 or 256;
[0050] S3.3 The final image semantic map representation is as follows:
[0051] ;
[0052] in, For node map, For the set of edges, Let be the set of edge features.
[0053] Further, the specific steps in S4 of calculating the node matching score and graph structure matching score of the image semantic graph and the text semantic graph through a cross-modal graph matching network, and fusing them to obtain the image-text similarity score, include:
[0054] S4.1, Define image node features Text node features Since the two have different dimensions, they need to be mapped to a unified form using linear projection.
[0055] ;
[0056] in, , These are the image projection weight matrix and the text projection weight matrix, respectively.
[0057] S4.2 Calculate the similarity matrix using the following formula:
[0058] ;
[0059] in, These are the projected image node features and text node features.
[0060] S4.3 To account for graph structure constraints, a more robust matching can be obtained through the optimal node alignment problem. Solving the optimal transport problem with marginal constraints yields the soft alignment matrix, as shown in the following formula:
[0061] ;
[0062] in, For the temperature parameter, Sinkhorn iteration ensures that the row and column are approximately 1;
[0063] S4.4, Node matching score can be defined as:
[0064] ;
[0065] Where N and M are the number of image nodes and the number of text nodes, respectively;
[0066] S4.5, Image Edge Features Text edge features Each projection is mapped to a unified space:
[0067] ;
[0068] in, Image edge features and text edge features; Projection weight matrices for image edges and text edges;
[0069] S4.6 Then calculate the edge similarity:
[0070] ;
[0071] S4.7 Calculate a weighted average of all edge matching scores based on node alignment weights:
[0072] ;
[0073] in, For the set of image edges, For the set of text edges, For image edges, For text edges, All are node alignment weights. To prevent division by zero of small constants;
[0074] S4.8, The final image-text similarity is:
[0075] .
[0076] Beneficial Effects: This invention provides a cross-modal remote sensing image-text retrieval method based on spatial layout awareness. By constructing image semantic graphs and text semantic graphs and introducing spatial layout constraints, it achieves fine-grained matching between image regions and text phrases. This effectively solves the problem of insufficient retrieval accuracy caused by neglecting spatial relationships of objects in existing technologies, significantly improving the accuracy of image-text retrieval involving spatial layout descriptions. Specifically, on the image side, this invention uses a target detection network to extract multiple ground object regions and generates node features for each region that fuse visual content and spatial location encoding. Simultaneously, it constructs edge features based on the relative position, distance, overlap, and scale differences between regions, forming a structured image semantic graph. On the text side, it parses entity phrases and their spatial relationship predicates through syntactic analysis and relation extraction, and uses a graph neural network to encode entity nodes and relation edges into a text semantic graph. This significantly improves the robustness of retrieval for complex spatial relationship descriptions. Attached Figure Description
[0077] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below:
[0078] Figure 1 This is a flowchart of a cross-modal remote sensing image retrieval method based on spatial layout awareness according to the present invention;
[0079] Figure 2 This is a flowchart of the image fine-grained feature extraction and spatial coding process of the present invention;
[0080] Figure 3 This is a flowchart of the text parsing and text semantic graph construction process of the present invention;
[0081] Figure 4 This is a flowchart of the cross-modal graph matching network of the present invention;
[0082] Figure 5 This is a table and graph showing the performance comparison of different models of this invention.
[0083] Figure 6 This is a diagram showing the experimental results of the model of this invention. Detailed Implementation
[0084] The embodiments of the present invention are described in detail below. Examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, but should not be construed as limiting the present invention.
[0085] like Figure 1 As shown, this invention provides a cross-modal remote sensing image and text retrieval method based on spatial layout awareness, comprising:
[0086] S1. Acquire the image, extract multiple target regions from the image, and generate region features and spatial location codes for each target region;
[0087] S2. Obtain the text query, parse the multiple entity phrases in the text query and the spatial relationships between the entity phrases, and construct a text semantic graph;
[0088] S3. Construct an image semantic graph by using the multiple target regions as nodes and the spatial relationships between regions as edges;
[0089] S4. Calculate the node matching score and graph structure matching score of the image semantic graph and the text semantic graph through a cross-modal graph matching network, and fuse them to obtain the image-text similarity score;
[0090] S5. Output the search results based on the image-text similarity score.
[0091] Furthermore, such as Figure 2 As shown, in step S1, an image is acquired, multiple target regions are extracted from the image, and region features and spatial location codes are generated for each target region. The specific steps are as follows:
[0092] S1.1 Input Image A target detection network is used to extract K candidate target regions. Each region i outputs a visual feature vector. Bounding box coordinates After normalization, the processing formula is as follows:
[0093] ;
[0094] Where W and H are the original image width and height; , These are the x-coordinates of the left and right boundaries of the bounding box of the i-th target region in the original image, respectively. , Let y be the upper and lower boundaries of the bounding box of the i-th target region in the original image;
[0095] S1.2 incorporates spatial information into the region representation, and modifies the normalized bounding box. Encoding is performed. Absolute position encoding is used to map the boundary coordinates into learnable embedding vectors, calculated as follows:
[0096] ;
[0097] MLP stands for Multilayer Perceptron. The dimension for location encoding;
[0098] S1.3 fuses visual features with location encoding to obtain the final region fusion feature. The calculation formula is as follows:
[0099] ;
[0100] [ ; ] indicates a splicing operation.
[0101] Furthermore, such as Figure 3 As shown, in step S2, the text query is obtained, multiple entity phrases in the text query and the spatial relationships between the entity phrases are parsed, and a text semantic graph is constructed. The specific steps are as follows:
[0102] S2.1 uses a pre-trained named entity recognition or noun block recognition model, assuming the extracted entity set is... Where P is the number of entities, for each entity phrase The semantic feature vector is extracted using a pre-trained language model (BERT), and the calculation formula is as follows:
[0103] ;
[0104] in, For the encoder of the language model, it can take CLS tags or average pooling output. For feature dimensions;
[0105] S2.2 Let the extracted spatial relation set be... , representing entities and There exists a spatial relationship between them. Each relationship can be represented as a triple. ,in For relational predicates, the relational predicates are mapped to learnable embedding vectors, and the calculation formula is as follows:
[0106] ;
[0107] in, For relational embedding functions;
[0108] Node S2.3 is Each node Let be the entity feature vector, and be the edge vector. To integrate node features into neighbor relationships and the global context, Graph Attention Networks (GAT) or Graph Convolutional Networks (GCN) can be used for message passing. The calculation formula is as follows:
[0109] ;
[0110] in, Let i be the set of neighboring nodes. Attention coefficient It is an aggregate function.
[0111] Furthermore, in step S3, the multiple target regions are used as nodes, and the spatial relationships between regions are used as edges to construct an image semantic graph. The specific steps are as follows:
[0112] S3.1 Using the normalized bounding box Calculate the following components:
[0113] Center coordinates: ;
[0114] Relative center offset: ;
[0115] Normalized central distance: ;
[0116] Intersection over Union (IoU): ;
[0117] Relative area ratio: ;
[0118] Direction angle: ;
[0119] The above components are concatenated to form the original geometric feature vector:
[0120] ;
[0121] in, Let x and y be the center coordinates of region i; Let x and y be the center coordinates of region j; The horizontal and vertical offsets are relative to the center. Let i be the bounding box of regions i and j; , The area of the overlapping region between the two frames, and the total area covered by the two frames; , Let i be the area of regions i and j; For the direction sine and cosine;
[0122] S3.2 uses a multilayer perceptron (MLP) to map the original geometric features into high-dimensional edge features. The calculation formula is as follows:
[0123] ;
[0124] in, Set it to 128 or 256;
[0125] The final image semantic map representation in S3.3 is as follows:
[0126] ;
[0127] in, For node map, For the set of edges, Let be the set of edge features.
[0128] Furthermore, such as Figure 4 As shown, in step S4, the node matching score and graph structure matching score of the image semantic graph and the text semantic graph are calculated through a cross-modal graph matching network, and then fused to obtain the image-text similarity score. The specific steps are as follows:
[0129] S4.1 Define image node features Text node features Since the two have different dimensions, they need to be mapped to a unified form using linear projection.
[0130] ;
[0131] in, , These are the image projection weight matrix and the text projection weight matrix, respectively.
[0132] S4.2 Calculate the similarity matrix using the following formula:
[0133] ;
[0134] in, These are the projected image node features and text node features.
[0135] S4.3 To consider graph structure constraints, a more robust matching can be obtained through the optimal node alignment problem. Solving the optimal transport problem with marginal constraints yields the soft alignment matrix, as shown in the following formula:
[0136] ;
[0137] in, For the temperature parameter, Sinkhorn iteration ensures that the row and column are approximately 1;
[0138] The S4.4 node matching score can be defined as:
[0139] ;
[0140] Where N and M are the number of image nodes and the number of text nodes, respectively;
[0141] S4.5 Image Edge Features Text edge features Each projection is mapped to a unified space:
[0142] ;
[0143] in, Image edge features and text edge features; Projection weight matrices for image edges and text edges;
[0144] S4.6 Then calculate the edge similarity:
[0145] ;
[0146] S4.7 calculates a weighted average of all edge matching scores based on node alignment weights:
[0147] ;
[0148] in, For the set of image edges, For the set of text edges, For image edges, For text edges, All are node alignment weights. To prevent division by zero of small constants;
[0149] The final image-text similarity score for S4.8 is:
[0150] .
[0151] like Figure 5 As shown in the model performance comparison table, the data is in percentage (%). The comparison models are described below:
[0152] RemoteCLIP: Based on the CLIP architecture, it performs continuous pre-training on remote sensing data, serving as the foundational model for cross-modal retrieval.
[0153] WSSCN: Proposes a complete semantic sparse coding network for constructing comprehensive and reliable feature representations.
[0154] SGPD: Transforms dense vectors into efficient sparse representations for retrieval using a sparsity-guided partially dense strategy.
[0155] Ours: We propose a cross-modal remote sensing image retrieval method based on spatial layout awareness.
[0156] This application employs two evaluation metrics, R@K and mR, to assess the performance of our model. R@K (Recall@K) represents the proportion of correct matches appearing in the top K search results across all query samples. R@1 (Recall@1) represents whether the correct result appears in the first (highest ranking) result; R@5 (Recall@5) represents whether the correct result appears in the top 5 results; R@10 (Recall@10) represents whether the correct result appears in the top 10 results. The mR metric represents the average of all R@K values, providing a more comprehensive assessment of the overall effectiveness of the model.
[0157] Our proposed method demonstrates comprehensive superiority across multiple metrics, particularly in image-to-text retrieval tasks, where its R@5 (56.37%) and R@10 (72.15%) are significantly higher than all comparable methods. Furthermore, EISCG achieves the best performance in mean recall (mR) at 56.13%, outperforming SGPD (53.78%), WSCN (51.25%), and RemoteCLIP (49.38%) by 2.35%, 4.88%, and 6.75%, respectively.
[0158] Figure 6 The qualitative results of our proposed method are presented. When a query is given, the correct matches of the top five retrieved results, from left to right, are highlighted in green, while some incorrect matches are marked in red. It can be seen that even for small objects such as vehicles and airplanes, our method is able to retrieve accurate matches in most cases. However, in some situations, some results marked "incorrect" may still correspond to semantically reasonable titles.
[0159] This invention presents a cross-modal remote sensing image-text retrieval method based on spatial layout awareness. By constructing image semantic graphs and text semantic graphs and introducing spatial layout constraints, it achieves fine-grained matching between image regions and text phrases, effectively solving the problem of insufficient retrieval accuracy caused by neglecting spatial relationships of objects in existing technologies, and significantly improving the accuracy of image-text retrieval involving spatial layout descriptions. Specifically, on the image side, this invention uses a target detection network to extract multiple ground object regions and generates node features for each region that fuse visual content and spatial location encoding. Simultaneously, it constructs edge features based on the relative position, distance, overlap, and scale differences between regions, forming a structured image semantic graph. On the text side, it parses entity phrases and their spatial relationship predicates through syntactic analysis and relation extraction, and uses a graph neural network to encode entity nodes and relation edges into a text semantic graph. This significantly improves the robustness of retrieval for complex spatial relationship descriptions.
[0160] The above-disclosed embodiments are merely one or more preferred embodiments of this application and should not be construed as limiting the scope of this application. Those skilled in the art can understand that all or part of the processes for implementing the above embodiments and equivalent changes made in accordance with the claims of this application still fall within the scope of this application.
Claims
1. A cross-modal remote sensing image and text retrieval method based on spatial layout awareness, characterized in that, include: S1. Acquire an image, extract multiple target regions from the image, and generate region features and spatial location codes for each target region. The spatial location codes are generated based on the bounding box information of the target region and fused with the region features of the corresponding target region to obtain image region node features. S2. Obtain the text query, parse the multiple entity phrases in the text query and the spatial relationship between the entity phrases, and construct a text semantic graph using the entity phrases as nodes and the spatial relationship between the entity phrases as edges. S3. Construct an image semantic graph by taking the multiple target regions as nodes and the spatial relationships between regions as edges, wherein the spatial relationships between regions are determined based on the geometric relationships between the target regions; S4. Calculate the node matching score and graph structure matching score of the image semantic graph and the text semantic graph using a cross-modal graph matching network, and fuse them to obtain a graph-text similarity score. Specifically, calculate the node similarity based on image region node features and text entity node features, and use the optimal transmission method to align image region nodes and text entity nodes to obtain a soft node alignment matrix A and node matching scores. Map the image edge features in the image semantic graph and the text edge features in the text semantic graph to a unified feature space, and calculate the edge relationship similarity between image edges and text edges. The image edge features are obtained based on the geometric relationship between corresponding target regions, and the text edge features are obtained based on the spatial relationship between corresponding entity phrases. For image edge (i,k) in the image semantic graph and text edge (j,l) in the text semantic graph, obtain the node alignment weights between image node i and text node j in the soft node alignment matrix A. And the node alignment weights between image node k and text node l The product of the alignment weights of the two nodes. The structure matching weights, which are used as the edge relationship similarity between the image edge (i,k) and the text edge (j,l), are used to aggregate the edge relationship similarity in the image semantic graph and the text semantic graph to obtain the graph structure matching score. S5. Output the search results based on the image-text similarity score.
2. The cross-modal remote sensing image and text retrieval method based on spatial layout awareness as described in claim 1, characterized in that, The specific steps of S1 include: S1.1 Input Image A target detection network is used to extract K candidate target regions, and each region i outputs a region visual feature vector. Bounding box coordinates After normalization, the processing formula is as follows: ; Where W and H are the original image width and height; , These are the x-coordinates of the left and right boundaries of the bounding box of the i-th target region in the original image, respectively. , Let y be the upper and lower boundaries of the bounding box of the i-th target region in the original image; S1.2 To incorporate spatial information into the region representation, the normalized bounding box... Encoding is performed using absolute positions, mapping boundary coordinates to position embedding vectors. The calculation formula is as follows: ; MLP stands for Multilayer Perceptron. The dimension for location encoding; S1.
3. Fuse visual features with location codes to obtain the final region fusion features. The calculation formula is as follows: ; [ ; ] indicates a splicing operation.
3. The cross-modal remote sensing image and text retrieval method based on spatial layout awareness as described in claim 1, characterized in that, The specific steps of S2 include: S2.1 Using a pre-trained named entity recognition or noun block recognition model, let the extracted entity set be... Where P is the number of entities, for each entity phrase The semantic feature vector is extracted using a pre-trained language model, and the calculation formula is as follows: ; in, For the encoder of the language model, take the CLS markers or average pooling output. For feature dimensions; S2.2, Let the extracted spatial relation set be... , representing entities and There exists a spatial relationship between them, and each relationship is represented as a triple. ,in For relational predicates, map the relational predicates to relational predicate embedding vectors. The calculation formula is as follows: ; in, For relational embedding functions; S2.3, Node is Each node Let be the entity feature vector, and be the edge vector. To integrate node features into neighbor relationships and the global context, a graph attention network or a graph convolutional network is used for message passing. The calculation formula is as follows: ; in, Let i be the set of neighboring nodes. Attention coefficient It is an aggregate function.
4. The cross-modal remote sensing image and text retrieval method based on spatial layout awareness as described in claim 1, characterized in that, The specific steps of S3 include: S3.1 Using the normalized bounding box Calculate the following components: Center coordinates: ; Relative center offset: ; Normalized central distance: ; Intersection over Union (IoU): ; Relative area ratio: ; Direction angle: ; The above components are concatenated to form the original geometric feature vector: ; in, Let x and y be the center coordinates of region i; Let x and y be the center coordinates of region j; The horizontal and vertical offsets are relative to the center. Let i be the bounding box of regions i and j; , The area of the overlapping region between the two frames, and the total area covered by the two frames; Let i be the area of regions i and j; For the direction sine and cosine; S3.
2. The original geometric features are mapped to high-dimensional edge features using a multilayer perceptron. The calculation formula is as follows: ; in, Set it to 128 or 256; S3.3 The final image semantic map representation is as follows: ; in, For node map, For the set of edges, Let be the set of edge features.
5. The cross-modal remote sensing image and text retrieval method based on spatial layout awareness as described in claim 1, characterized in that, The specific steps of S4 include: S4.1, Define image node features Text node features Since the two have different dimensions, they need to be mapped to a unified form using linear projection. ; in, , These are the image projection weight matrix and the text projection weight matrix, respectively. S4.2 Calculate the similarity matrix using the following formula: ; in, These are the projected image node features and text node features. S4.3 To account for graph structure constraints, the optimal transport problem with marginal constraints is solved by optimal node alignment, yielding the soft alignment matrix, as shown in the following formula: ; in, For the temperature parameter, the Sinkhorn iteration is used to normalize the rows and columns of the matrix; S4.4, The node matching score is defined as: ; Where N and M are the number of image nodes and the number of text nodes, respectively; S4.5, Image Edge Features Text edge features Each projection is mapped to a unified space: ; in, Image edge features and text edge features; Projection weight matrices for image edges and text edges; S4.6 Then calculate the edge similarity: ; S4.7 Calculate a weighted average of all edge matching scores based on node alignment weights: ; in, For the set of image edges, For the set of text edges, For image edges, For text edges, All are node alignment weights. To prevent small constants from being divided by zero; S4.8, The final image-text similarity is: 。
Citation Information
Patent Citations
Image-text retrieval method and system based on attention mechanism and gating mechanism
CN112966135A
Remote sensing image cross-modal retrieval method based on layout semantic joint significant representation
CN116561365A
Cross-modal remote sensing image-text retrieval method based on multistage semantic collaborative matching
CN120336574A
Multi-modal semantic and physical law driven remote sensing image generation method
CN121353446A