Cross-modal remote sensing image-text retrieval method based on multistage semantic collaborative matching

Through the multi-level semantic collaborative matching method, the global, regional and pixel-level features of remote sensing images are extracted, which solves the problem of insufficient understanding of global and fine-grained features in remote sensing graphics and text retrieval, and achieves higher precision graphics and text matching.

CN120336574AActive Publication Date: 2025-07-18CHINA UNIV OF MINING & TECH +1

Patent Information

Application Number
CN202510813500.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-07-18
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

The existing remote sensing graphic search methods cannot effectively understand the global and fine-grained characteristics of remote sensing images, resulting in poor robustness in handling fine-grained problems.

Method used

The multi-level semantic collaborative matching method is used to extract the region of interest through the image preprocessing module, combine the multi-level visual attention mechanism and the Transformer model to extract the global, regional and pixel-level features of the image, and use the cross attention mechanism and dynamic fusion mechanism to calculate the similarity between the remote sensing image and the text description.

Benefits of technology

It improves the semantic consistency between image and text features, enhances the accuracy of picture and text matching, can effectively capture tiny features such as land objects edges and textures in remote sensing images, and improves the retrieval fineness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336574A_ABST
    Figure CN120336574A_ABST
Patent Text Reader

Abstract

The invention provides a cross-modal remote sensing image-text retrieval method based on multistage semantic collaborative matching, which comprises the following steps of: extracting a region of interest by using a semantic segmentation algorithm through an image preprocessing module, segmenting a remote sensing image into a plurality of regions and generating image blocks; respectively extracting global features, regional features and pixel-level features of the image, and carrying out fine-grained coding on key regions such as fine-grained ground feature edges and the like; the text multi-level coding module is used for carrying out three-level feature coding of documents, sentences and words on texts based on a pre-training language model to ensure multi-level understanding of the texts; in the multi-level matching and fusion module, the similarity between the remote sensing image and the text description is calculated through a cross attention mechanism, weighted fusion is carried out on features of all levels, and finally a retrieval score is output. The method not only improves the accuracy and robustness of image-text retrieval, but also can be widely applied to the fields of remote sensing monitoring, environment change recognition, geographic information systems and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image processing, and specifically to a cross-modal remote sensing image-text retrieval method based on multi-level semantic collaborative matching. Background Art

[0002] Remote sensing image-text retrieval is a multi-modal information retrieval technology that combines remote sensing images and natural language text descriptions. With the development of remote sensing technology, the application fields of remote sensing images are becoming more and more extensive, covering multiple aspects such as environmental monitoring, urban planning, disaster warning, and agricultural assessment. However, the high dimensionality, complexity, and professionalism of remote sensing images make traditional image retrieval methods face challenges in practical applications. Therefore, how to effectively extract useful information from remote sensing images and match it with relevant text descriptions has become a key issue.

[0003] The significance of remote sensing image-text retrieval lies in its ability to achieve efficient retrieval of remote sensing data through semantic associations between text and images. For example, through natural language queries input by users, relevant images can be quickly found from a large number of remote sensing images, greatly improving the efficiency of information retrieval. In addition, remote sensing image-text retrieval can also support cross-modal learning. Even if the feature expressions of remote sensing images and texts are different, the system can still achieve comprehensive understanding of different modal data through effective feature alignment and matching.

[0004] Many existing methods rely on simple keyword-based matching or rough feature similarity calculation and cannot fully understand the deep semantics in images and texts. For example, traditional methods based on convolutional networks, due to the limitation of the receptive field, often can only capture information in local regions and ignore the global information and long-range dependencies in remote sensing images. At the same time, existing methods usually focus on macroscopic-level image and text features and are insufficient in processing fine-grained features in remote sensing images, which makes the retrieval results less robust when dealing with fine-grained problems. Therefore, developing a practical and efficient remote sensing image-text retrieval method has become an urgent need. Summary of the Invention

[0005] The purpose of the present invention is to provide a cross-modal remote sensing image-text retrieval method based on multi-level semantic collaborative matching.

[0006] To achieve the above purpose, the present invention provides a cross-modal remote sensing image-text retrieval method based on multi-level semantic collaborative matching, including the following steps: S1. Use a semantic segmentation algorithm through an image preprocessing module to extract regions of interest, divide the remote sensing image into multiple regions, and generate image patches; S2. Use a hybrid architecture that combines a convolutional neural network and a Transformer model with a multi-level visual attention mechanism to extract the global features, regional features, and pixel-level features of the image in the multi-level feature extraction module, and perform fine-grained encoding on key regions such as fine-grained object edges; S3. Use a text multi-level encoding module to perform feature encoding on the text at the document, sentence, and word levels based on a pre-trained language model to improve the model's multi-level understanding of the text; S4. In the multi-level matching and fusion module, calculate the similarity between the remote sensing image and the text description through a cross-attention mechanism and a dynamic fusion mechanism, and perform weighted fusion on the features at all levels; S5. Based on the multi-level feature matching and fusion, output the score of the image-text retrieval.

[0007] Further, step S1 is specifically as follows: S1.1-1. Use a filter-based denoising method (Gaussian filter) to denoise the image. Gaussian filter is a commonly used image smoothing technique for removing noise in the image while retaining the edge information of the image. Let the original input image be I(x, y), the Gaussian filter be G(x, y), and the mathematical expression of the two-dimensional Gaussian function is: ; where G(x, y) is the value of the Gaussian function at the position (x, y), σ is the standard deviation of the Gaussian function, and x, y are the coordinates of the pixel, the offset relative to the current pixel; S1.1-2. Perform edge padding at the boundary of the input image I(x, y), and then perform a convolution operation on each pixel with the Gaussian kernel to generate a filtered image , and calculate the output value of each pixel point The formula is as follows: ; where I(x-i, y-i) is the pixel value of the original image at the position (x-i, y-i); G(x, y) is the weight at the position (x, y) in the Gaussian kernel, and k represents the radius size of the kernel; S1.1-3. The Gaussian kernel needs to be normalized so that the sum of all weights is equal to 1. The normalization formula is as follows: ; where G(x, y) is the weight at the position (x, y) in the Gaussian kernel, k represents the radius size of the kernel, and x, y are the coordinates of the pixel; S1.2-1. Extract the region of interest using the semantic segmentation algorithm. The semantic segmentation algorithm refers to the use of computer vision technology to assign specific semantic labels to each pixel in the image, achieving precise segmentation of different regions or objects in the image, dividing the remote sensing image into multiple regions and generating image patches. The network extracts features from the image through convolutional operations, extracting low-level (such as edges, textures, etc.) and high-level (such as objects, regions, etc.) features. The formula for the convolutional operation is: ; where p is the input image, is the convolutional kernel, and f is the output after convolution; S1.2-2. The pooling operation is used to reduce the spatial resolution of the image, reduce the computational amount, and retain features at the same time. The formula for the pooling operation is: ; where R is the pooling window, p is the input image, and Z is the output after pooling; S1.2-3. After multiple convolutions and poolings, the resolution of the image will decrease. Therefore, it is necessary to restore the spatial resolution of the image through upsampling (transposed convolution). The formula for the transposed convolution operation is: ; where I(x,y) is the pixel of the image after upsampling, and G(i,j) is the transposed convolution kernel; S1.2-4. The output of each pixel is the probability that it belongs to each category. Usually, the softmax function is used to convert these scores into probability values. The formula for the softmax function is: ; where, is the score of pixel belonging to , is the probability that pixel belongs to category , and C is the total number of categories; S1.2-5. After classification, each pixel of the image will be assigned a category label. According to the category label, the region of a specific category can be extracted to form the region of interest AOI. The formula for extracting AOI is: ; where, AOI is the region of interest, is the probability that the pixel (x,y) belongs to category , and threshold is the set threshold; Furthermore, step S2 is specifically: S2.1-1. In global feature extraction, convolutional layers with relatively large convolutional kernels (such as 5*5 or 7*7) are usually used to process images. This can help the network capture context information in a larger range, thereby obtaining a global understanding of the image. The formula for the convolution operation is as shown in S1.2-1; S2.1-2. When performing global feature extraction, the Transformer model utilizes the self-attention mechanism to model the global dependencies in the image. The self-attention mechanism is a core component of the Transformer model and is used to calculate the correlations between elements within the input sequence, helping the model capture the dependencies between elements at different positions in the sequence, regardless of how far apart they are. Suppose an input feature matrix , where , is the set of real numbers. The attention output is calculated through queries (Q), keys (K), and values (V), and the formula is as follows: ; ; Among them, Q, K, V, and T are the query, key, value matrices, and transpose respectively, is the dimension of the key vector, A is the calculated attention weight matrix, is the weighted global feature output, and softmax is the normalized exponential function. Its core role is to convert a set of input numerical values into a probability distribution; S2.1-3. Calculate the attention weights through the visual attention mechanism and focus the attention on the key regions of the image. The formula is as follows: ; Among them, is the weighted global feature output, is the global attention weight, and softmax is the normalized exponential function. Its core role is to convert a set of input numerical values into a probability distribution; S2.2-1. In local feature extraction, relatively small convolutional kernels (such as 3*3 or 5*5) are usually used to extract features in the image through the sliding window method. Through local feature extraction, the network can focus on the features of different parts of the image. The formula for the convolution operation is as shown in S1.2-1; S2.2-2. Unify the size of the regions through pooling operations and provide a region feature map of a fixed size for subsequent processing. Then, calculate the weighted output of the local features. The formula is as follows: ; Among them, m is the input region feature map, is regional pooling, is max pooling, is average pooling; S2.2-3. The network can adaptively focus attention on relevant regions of the image, especially regions with key semantic information. The cross-attention mechanism formula is as follows: ; Among them, is the regional feature extracted by the network, is the regional attention weight. Softmax is a normalized exponential function, and its core function is to convert a set of input numerical values into a probability distribution; S2.3-1. In pixel-level feature extraction, it usually focuses on fine-grained features of the image, such as object edges, detailed textures, etc. The formula for the convolution operation is as shown in S1.2-1. The Sobel operator is used to calculate the gradient of the image and extract edges. Its convolution kernel and the calculation of the gradient intensity of the image are as follows: ; ; ; S2.3-2. In pixel-level feature extraction, the cross-attention mechanism assigns weights to each pixel. The formula is as follows: ; Among them, is the pixel feature extracted by the network, is the pixel attention weight. Softmax is a normalized exponential function, and its core function is to convert a set of input numerical values into a probability distribution.

[0008] Furthermore, step S3 is specifically as follows: S3.1-1. Assume that the input text T is segmented into document-level (M sentences), and each sentence is further tokenized into . Using BERT as the basic encoder, word, sentence, and document-level features are generated as pixel features extracted by the network. It realizes in-depth semantic understanding of the text through a bidirectional self-attention mechanism, significantly improving the performance of multiple natural language processing tasks. The word-level feature formula is as follows: ; Among them, is the set of real numbers, represents a real number matrix space with a dimension of , L is the total length of the sequence, and d is the BERT hidden layer dimension (such as 768); S3.1-2. For the word-level feature of each sentence Average pooling is performed, and the formula is as follows: ; ; where N is the number of word segments in each sentence, is the feature of the sentence, is the set of all sentence features; S3.1-3. Self-attention aggregation is used for the sentence-level feature , and the output is the global document vector , and the formula is as follows: ; Furthermore, step S4 is specifically as follows: The dynamic fusion mechanism is a strategy for multi-modal or multi-level feature fusion. Its core idea is to adaptively allocate weights between different features or information sources according to the specific input and task requirements to achieve more flexible and effective feature integration, and perform weighted summation on document-level, sentence-level, and word-level features, and weight the contribution of each layer through the learned weights. The weighted fusion formula is: ; where, is the document-level feature, is the sentence feature, is the word-level feature.

[0009] Beneficial effects: Through multi-level feature extraction and alignment, the present invention can capture the global semantics and local detail information of images and texts. At the same time, by using the semantic collaborative matching mechanism, the semantic consistency between image and text features is improved, the accuracy of image-text matching is enhanced, and through fine-grained coding, it can effectively capture tiny features such as object edges and textures in remote sensing images, improving the fineness of retrieval. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 is the flow chart of the present invention; Figure 2 The schematic diagram of the principle of the Gaussian filter adopted by the preprocessing module of the present invention; Figure 3 is the schematic diagram of the principle of the multi-level visual attention mechanism adopted by the present invention; Figure 4 is the flow chart of the Transformer model adopted in the multi-level feature extraction module of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0011] The present invention will be further described below with reference to the accompanying drawings.

[0012] As Figure 1As shown in the figure, a cross-modal remote sensing image-text retrieval method based on multi-level semantic collaborative matching includes the following steps: S1. Use a semantic segmentation algorithm through an image preprocessing module to extract regions of interest, segment the remote sensing image into multiple regions, and generate image patches. S2. In a multi-level feature extraction module, adopt a hybrid architecture combining a convolutional neural network and a Transformer model with a multi-level visual attention mechanism to extract the global feature, regional feature, and pixel-level feature of the image respectively, and perform fine-grained encoding on key regions such as fine-grained ground object edges. S3. Use a text multi-level encoding module to perform three-level feature encoding of the text at the document, sentence, and word levels based on a pre-trained language model to improve the model's multi-level understanding of the text. S4. In a multi-level matching and fusion module, calculate the similarity between the remote sensing image and the text description through a cross-attention mechanism and a dynamic fusion mechanism, and perform weighted fusion on each level of features. S5. Based on the multi-level feature matching and fusion, output the score of the image-text retrieval. As a preferred implementation, step S1 is specifically as follows: S1.1-1. Use a filter-based denoising method (Gaussian filter) to denoise the image. As Figure 1 shown, Gaussian filtering is a common image smoothing technique used to remove noise in the image while retaining the edge information of the image. Let the original input image be I(x, y), the Gaussian filter be G(x, y), and the mathematical expression of the two-dimensional Gaussian function is: ; where G(x, y) is the value of the Gaussian function at the position (x, y), σ is the standard deviation of the Gaussian function, and x, y are the coordinates of the pixel, the offset relative to the current pixel. S1.1-2. Perform edge padding at the boundary of the input image I(x, y), and then perform a convolution operation on each pixel point with the Gaussian kernel to generate a filtered image , and calculate the output value of each pixel point. The formula is as follows: ; where I(x - i, y - i) is the pixel value of the original image at the position (x - i, y - i); G(x, y) is the weight at the position (x, y) in the Gaussian kernel, and k represents the radius size of the kernel. S1.1-3. The Gaussian kernel needs to be normalized so that the sum of all weights is equal to 1. The normalization formula is as follows: ; Among them, G(x,y) is the weight at position (x,y) in the Gaussian kernel, k represents the radius size of the kernel, and x,y are the coordinates of the pixels; S1.2-1. Use the semantic segmentation algorithm to extract the region of interest. The semantic segmentation algorithm refers to using computer vision technology to assign specific semantic labels to each pixel in the image, achieving precise segmentation of different regions or objects in the image, segmenting the remote sensing image into multiple regions and generating image patches. The network extracts features from the image through convolutional operations, extracting low-level (such as edges, textures, etc.) and high-level (such as objects, regions, etc.) features. The formula for the convolutional operation is: ; Among them, p is the input image, is the convolutional kernel, and f is the output after convolution; S1.2-2. The pooling operation is used to reduce the spatial resolution of the image, reduce the computational amount, and retain features at the same time. The formula for the pooling operation is: ; Among them, R is the pooling window, p is the input image, and Z is the output after pooling; S1.2-3. After multiple convolutions and poolings, the resolution of the image will decrease. Therefore, it is necessary to restore the spatial resolution of the image through upsampling (deconvolution). The formula for the deconvolution operation is: ; Among them, I(x,y) is the pixel of the image after upsampling, and G(i,j) is the deconvolution kernel; S1.2-4. The output of each pixel is the probability that it belongs to each category. Usually, the softmax function is used to convert these scores into probability values. The formula for the softmax function is: ; Among them, is the score of the pixel belonging to , is the probability that the pixel belongs to the category , and C is the total number of categories; S1.2-5. After classification, each pixel of the image will be assigned a category label. According to the category label, the region of a specific category can be extracted to form the region of interest AOI. The formula for extracting AOI is: ; Among them, AOI is the region of interest, is the probability that the pixel (x,y) belongs to the category , and threshold is the set threshold; As a preferred implementation, step S2 is specifically as follows: S2.1-1. In global feature extraction, a convolutional layer with a relatively large convolutional kernel (such as 5*5 or 7*7) is usually used to process the image. This can help the network capture a larger range of context information, thereby obtaining a global understanding of the image. The formula for the convolutional operation is as shown in S1.2-1; S2.1-2. When performing global feature extraction, the Transformer model uses the self-attention mechanism to model the global dependencies in the image, as Figure 4 shown. The self-attention mechanism is a core component of the Transformer model and is used to calculate the correlations between elements within the input sequence, helping the model capture the dependencies between elements at different positions in the sequence, regardless of how far apart they are, as Figure 3 shown. Suppose an input feature matrix is input, where , is the set of real numbers. The attention output is calculated through queries (Q), keys (K), and values (V), and the formula is as follows: ; ; where Q, K, V, and T are the query, key, value matrices, and transpose respectively, is the dimension of the key vector, A is the calculated attention weight matrix, is the weighted global feature output, and softmax is the normalized exponential function whose core role is to convert a set of input numerical values into a probability distribution; S2.1-3. Calculate the attention weights through the visual attention mechanism and focus the attention on the key regions of the image. The formula is as follows: ; where, is the weighted global feature output, is the global attention weight, and softmax is the normalized exponential function whose core role is to convert a set of input numerical values into a probability distribution; S2.2-1. In region feature extraction, a relatively small convolutional kernel (such as 3*3 or 5*5) is usually used to extract features in the image through the sliding window method. Through region feature extraction, the network can focus on the features of different parts of the image. The formula for the convolutional operation is as shown in S1.2-1; S2.2-2. Unify the size of the regions through pooling operations and provide a region feature map of a fixed size for subsequent processing, and then calculate the weighted output of the region features. The formula is as follows: ; Among them, is the input regional feature map, is regional pooling, is max pooling, is average pooling; S2.2-3. The network can adaptively focus attention on relevant regions of the image, especially regions with key semantic information. The cross-attention mechanism formula is as follows: ; Among them, is the regional feature extracted by the network, is the regional attention weight, and softmax is the normalized exponential function. Its core role is to convert a set of input numerical values into a probability distribution; S2.3-1. In pixel-level feature extraction, it usually focuses on fine-grained features of the image, such as object edges, detailed textures, etc. The formula for the convolution operation is as shown in S1.2-1. The Sobel operator is used to calculate the gradient of the image and extract edges. Its convolution kernel and the calculation of the gradient intensity of the image are as follows: ; ; ; S2.3-2. In pixel-level feature extraction, the cross-attention mechanism assigns weights to each pixel. The formula is as follows: ; Among them, is the pixel feature extracted by the network, is the pixel attention weight, and softmax is the normalized exponential function. Its core role is to convert a set of input numerical values into a probability distribution.

[0013] As a preferred implementation manner, step S3 is specifically as follows: S3.1-1. Assume that the input text T is segmented into document-level, (M sentences), and each sentence is further segmented into . Using BERT as the basic encoder, the word, sentence, and document-level features are pixel features extracted through the network. It realizes deep semantic understanding of the text through the bidirectional self-attention mechanism, significantly improving the performance of multiple natural language processing tasks. The word-level feature formula is as follows: ; Among them, is the set of real numbers, represents a dimension of The real matrix space, L is the total length of the sequence, and d is the dimension of the BERT hidden layer (e.g., 768); S3.1-2. For each sentence at the word level perform average pooling, and the formula is as follows: ; ; where N is the number of word segments in each sentence, is the feature of the sentence, is the set of all sentence features; S3.1-3. For the sentence-level feature use self-attention aggregation, and the output is the global document vector , and the formula is as follows: ; As a preferred implementation, step S4 is specifically: The dynamic fusion mechanism is a strategy for multi-modal or multi-level feature fusion. Its core idea is to adaptively allocate weights between different features or information sources according to specific inputs and task requirements to achieve more flexible and effective feature integration. Perform weighted summation on document-level, sentence-level, and word-level features, and weight the contributions of each layer through learned weights. The weighted fusion formula is: ; where, is the document-level feature, is the sentence feature, is the word-level feature.

Claims

1. A cross-modal remote sensing image-text retrieval method based on multi-level semantic collaborative matching, characterized in that It includes the following steps: S1. Use the semantic segmentation algorithm in the image preprocessing module to extract the region of interest, segment the remote sensing image into multiple regions and generate image patches; S2. Adopt a hybrid architecture combining a convolutional neural network and a Transformer model with a multi-level visual attention mechanism to extract the global features, regional features and pixel-level features of the image respectively in the multi-level feature extraction module, and perform fine-grained encoding on the fine-grained key regions of the object edges; S3. Adopt the text multi-level encoding module to perform feature encoding on the text at the document, sentence and word levels based on the pre-trained language model to improve the model's multi-level understanding of the text; S4. In the multi-level matching and fusion module, calculate the similarity between the remote sensing image and the text description through the cross-attention mechanism and the dynamic fusion mechanism, and perform weighted fusion on the features at all levels; S5. Based on the multi-level feature matching and fusion, output the score of the image-text retrieval.

2. The cross-modal remote sensing image-text retrieval method based on multi-level semantic collaborative matching according to claim 1, wherein Step S1 is specifically as follows: S1.1-1. Use a filter-based denoising method to denoise the image. Let the original input image be I(x, y), the Gaussian filter be G(x, y), and the mathematical expression of the two-dimensional Gaussian function be: ; where G(x,y) is the value of the Gaussian function at the position (x,y), σ is the standard deviation of the Gaussian function, and x,y are the coordinates of the pixels; S1.1-2. Perform edge padding at the boundary of the input image I(x, y), and then perform a convolution operation on each pixel with a Gaussian kernel to generate a filtered image , and calculate the output value of each pixel . The formula is as follows: ; where I(x-i, y-i) is the pixel value of the original image at the position (x-i, y-i); G(x,y) is the weight at the position (x,y) in the Gaussian kernel, and k represents the radius size of the kernel; S1.1-3. The Gaussian kernel needs to be normalized so that the sum of all weights is equal to 1. The normalization formula is as follows: ; where G(x,y) is the weight at the position (x,y) in the Gaussian kernel, k represents the radius size of the kernel, and x,y are the coordinates of the pixels; S1.2-1. Use the semantic segmentation algorithm to extract the region of interest, segment the remote sensing image into multiple regions and generate image patches. The network extracts features from the image through convolutional operations, extracting low-level features and high-level features. The formula for the convolutional operation is: ; where p is the input image, is the convolution kernel, and f is the output after convolution; S1.2-2. Pooling operations are used to reduce the spatial resolution of the image, reduce the computational amount, and retain features at the same time. The formula for the pooling operation is: ; where R is the pooling window, p is the input image, and Z is the output after pooling; S1.2-3. After multiple convolutions and poolings, the resolution of the image will be reduced. The spatial resolution of the image is restored through upsampling. The formula for the upsampling operation is: ; where I(x,y) is the pixel of the image after upsampling, and G(i,j) is the deconvolution kernel; S1.2-4. The output of each pixel is the probability that it belongs to each category. Use the softmax function to convert these scores into probability values. The formula for the softmax function is: ; Among them, is the pixel belongs to score, is the pixel belongs to the category probability, and C is the total number of categories; S1.2-5. After classification, each pixel of the image will be assigned a category label. According to the category label, extract the regions of specific categories to form the region of interest. The formula for extracting the region of interest is: ; Among them, AOI is the region of interest, is the probability that the pixel (x, y) belongs to the category , and threshold is the set threshold.

3. The cross-modal remote sensing image-text retrieval method based on multi-level semantic collaborative matching according to claim 2, wherein, Step S2 is specifically as follows: S2.1-1. In global feature extraction, a convolutional layer with a larger convolutional kernel is used to process the image, and the formula for the convolutional operation is as shown in step S1.2-1; S2.1-2. When extracting global features, the Transformer model uses the self-attention mechanism to model the global dependencies in the image. Assume that a feature matrix is input , where , is the set of real numbers. The attention output is calculated through queries, keys, and values, and the formula is as follows: ; ; where Q, K, V, and T are the query, key, value matrices, and transpose respectively, is the dimension of the key vector, A is the calculated attention weight matrix, is the output of the weighted global feature, and softmax is the normalized exponential function whose core role is to convert a set of input values into a probability distribution; S2.1-3. Calculate the attention weights through the visual attention mechanism, and focus the attention on the key regions of the image. The formula is as follows: ; Among them, is the output of the weighted global feature, is the global attention weight; S2.2-1. In region feature extraction, a smaller convolutional kernel is usually used to extract features in the image through the sliding window method. Through region feature extraction, the network can focus on the features of different parts of the image. The formula for the convolutional operation is as shown in S1.2-1; S2.2-2. Uniform the size of the region by performing a pooling operation on the region of interest, and provide a region feature map of a fixed size for subsequent processing. Then calculate the weighted output of the region features. The formula is as follows: ; Among them, m is the input regional feature map, is regional pooling, is max pooling, is average pooling; S2.2-3. The network can adaptively focus the attention on the relevant regions of the image, especially the regions with key semantic information. The formula for the cross-attention mechanism is as follows: ; Among them, is the regional feature extracted through the network, is the regional attention weight; S2.3-1. In pixel-level feature extraction, usually focus on the fine-grained features of the image. The formula for the convolutional operation is as shown in S1.2-1. The Sobel edge detection operator is used to calculate the gradient of the image and extract the edges. Its convolutional kernel and the calculation of the gradient intensity of the image are as follows: ; ; ; S2.3-2. In pixel-level feature extraction, the cross-attention mechanism assigns weights to each pixel. The formula is as follows: ; Among them, is the pixel feature extracted through the network, is the pixel attention weight.

4. The cross-modal remote sensing image-text retrieval method based on multi-level semantic collaborative matching according to claim 1, characterized in that Step S3 is specifically as follows: S3.1-1. Assume that the input text T is segmented at the document level , that is, into M sentences, and each sentence is further tokenized into . Using BERT as the base encoder, pixel features extracted through the network for word, sentence, and document-level features are generated. The formula for word-level features is as follows: ; Among them, is the set of real numbers, represents a real matrix space of dimension , where L is the total length of the sequence and d is the dimension of the BERT hidden layer; S3.1-2. Average pool the word-level features of each sentence The word-level features are averaged using the following formula: ; ; Among them, N is the number of word segments in each sentence, which is the feature of the sentence, and is the set of all sentence features; S3.1-3. Aggregate the sentence-level features Use self-attention aggregation, and the output is the global document vector , and the formula is as follows: 。 5. The cross-modal remote sensing image-text retrieval method based on multi-level semantic collaborative matching according to claim 1, wherein Step S4 is specifically as follows: Perform a weighted sum on the document-level, sentence-level, and word-level features, and weight the contributions of each layer through the learned weights. The weighted fusion formula is: ; Among them, is a document-level feature, is the set of all sentence features, is a word-level feature.

Citation Information

Patent Citations

  • Abstract generation method and apparatus, and computer device

    CN108280112A

  • Semantic segmentation method for large-format remote sensing image from PATCH to REGION architecture

    CN116310325A

  • Extracting mentions of complex relation types from documents

    US20220284192A1

Cited By

  • Cross-modal representation learning and retrieval method and system for grain production

    CN120705355A

  • Tumor HER2 expression grading method based on HE dyeing image

    CN120808885A

  • Multi-core feature representation learning system and method for enhancing remote sensing image based on content retrieval

    CN120853029A

  • Remote sensing image-text retrieval method based on knowledge enhancement and asymmetric structure

    CN120973970A

  • Remote sensing image-text retrieval method based on knowledge enhancement and asymmetric structure

    CN120973970B