Text-guided Visual Localization Method for Remote Sensing Images

By constructing a remote sensing image visual positioning network model of the parallel arranged text-guided visual feature generation network and text encoder, the problem of low visual positioning accuracy of remote sensing images is solved, and higher positioning accuracy and more accurate target recognition are achieved.

CN116958829BActive Publication Date: 2025-07-01XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310866853.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-14
Publication Date
2025-07-01
Estimated Expiration
2043-07-14

AI Technical Summary

Technical Problem

The existing remote sensing image visual positioning methods have challenges in improving positioning accuracy, especially because the remote sensing image is large in size and insignificant object characteristics, resulting in low accuracy of visual positioning.

Method used

A remote sensing image visual positioning network model including parallel arrangement of text-guided visual feature generation network and text encoder is constructed, and the generation of visual features is guided at the channel level and spatial level through global text features, making full use of the semantic and spatial information of text features to improve the quality of visual features.

Benefits of technology

The accuracy of visual positioning of remote sensing images is significantly improved. By fully utilizing the semantic and spatial information of text features, the ambiguity in the semantic information is reduced, and the target features that are not significant enough in the original feature map of the remote sensing image are effectively supplemented.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116958829B_ABST
    Figure CN116958829B_ABST
Patent Text Reader

Abstract

The invention proposes a remote sensing image visual positioning method based on text guidance, and the implementation steps are as follows: obtaining a training sample set and a test sample set; constructing a remote sensing image visual positioning network model: comprising a text-guided visual feature generation network, a text encoder, a multimodal fusion network and a positioning network; initializing parameters; training the visual positioning network model; updating the parameters of the visual positioning network model; and obtaining a visual positioning detection result. The positioning network model constructed by the invention uses global text features to guide the generation of visual features at the channel level and the spatial level, makes full use of the global semantic information of the text features, reduces the ambiguity in the semantic information, and uses text features of different levels to guide visual features of different scales in multiple stages, makes full use of the shallow features and deep features of the text, and the spatial information of visual feature maps of different scales, supplements the target features that are not prominent enough in the original feature map, and effectively improves the accuracy of remote sensing image visual positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and relates to a text-guided visual positioning method for remote sensing images, which can be applied to fields such as environmental protection and disaster emergency. Background Art

[0002] Remote sensing images are images that record the electromagnetic waves of various ground objects detected by remote sensors under the conditions of being far from the target and non-contact with the target object. The remote sensing image positioning method is a method for positioning and identifying targets in remote sensing images, which can effectively use the ground object information in remote sensing images for production and life. Currently, it has been widely applied in fields such as environmental protection, disaster emergency, urban planning, agricultural production, geological disaster investigation and treatment, and earth resource investigation. The goal of the remote sensing image visual positioning method is to locate the specified target in the remote sensing image according to the user's text description, and its key lies in improving the positioning accuracy. However, due to the characteristics of large scale and insignificant object features of remote sensing images, achieving accurate target positioning is still a huge challenge.

[0003] In order to improve the accuracy of remote sensing image visual positioning, prior art has made explorations. For example, Yang Zhan designed a multi-scale cross-modal fusion method based on Transformer in the paper "RSVG: Exploring Data and Models for Visual Grounding on Remote Sensing Data" published in the journal IEEE Transactions on Geoscience and Remote Sensing on February 28, 2023. The visual branch uses ResNet50 to extract multi-scale visual features, and the visual features of different resolutions are simply concatenated as multi-scale visual features; the text branch uses Bert to extract text features, and then the [CLS] embedding and word features are concatenated as the final text features; after obtaining the features of the two modalities, they are input into a multi-stage Transformer decoder for feature fusion, and finally target positioning is performed according to the fused features. This method takes into account the large scale characteristic of remote sensing images, makes relatively full use of the multi-scale features extracted by the backbone network, and uses text features for guidance, so that it can combine the effective information from multi-level and multi-modal features, and to a certain extent improves the accuracy of remote sensing image visual positioning. However, since it only uses the concatenation operation for the [CLS] embedding and word features, the semantic and spatial information of the text features is still not fully utilized, and the weight ratio of different scale visual features is not fully considered, resulting in a low accuracy rate of its visual positioning. Summary of the Invention

[0004] The object of the present invention is to overcome the defects of the above-mentioned existing technologies, and a semantic-guided visual positioning method for remote sensing images is proposed, aiming to improve the positioning accuracy of the visual positioning method for remote sensing images.

[0005] To achieve the above object, the technical solution adopted by the present invention includes the following steps:

[0006] (1) Obtain a training sample set and a test sample set:

[0007] Label the targets contained in each of the K remote sensing images obtained, and form a training sample set R1 with M remote sensing images and their corresponding labels and texts, and form a test sample set E1 with the remaining K - M remote sensing images and their corresponding labels and texts, where K≥500.

[0008] (2) Construct a visual positioning network model G for remote sensing images:

[0009] Construct a visual positioning network model G for remote sensing images including a text-guided visual feature generation network and a text encoder arranged in parallel, and a multi-modal fusion network and a positioning network cascaded in sequence at the output ends of the text-guided visual feature generation network and the text encoder; the text-guided visual feature generation network includes a downsampling block cascaded in sequence, N composite modules composed of cascaded visual feature extraction modules and text-guided fusion modules, and R Transformer encoders; the text encoder includes N text feature extraction blocks cascaded in sequence, and the nth text feature extraction block is also connected to the corresponding nth text-guided fusion module; the multi-modal fusion network includes a language-guided module and a context-guided module arranged in parallel, and S Transformer decoders cascaded in sequence at the output ends of the language-guided module and the context-guided module; where N≥1, R≥1, S≥1.

[0010] (3) Initialize the parameters:

[0011] Initialize the number of iterations as h, the maximum number of iterations as H, H≥150, and the weights and bias parameters of the visual positioning network model G at the hth iteration are w h and b h respectively, and let h = 0, G h = G; h

[0012] (4) Train the visual positioning network model G:

[0013] Randomly and with replacement select L training samples from the training sample set R1 as the input of the visual positioning network model G for forward propagation to obtain L visual positioning results, where 1≤L≤M.

[0014] (5) Update the parameters of the visual localization network model:

[0015] Based on the L visual localization results obtained in step (4), update the weights and bias parameters w h and b h of the visual localization network model G h to obtain the network model G of this iteration h ; and determine whether h > H holds. If so, obtain the trained visual localization network model G*. Otherwise, set G = G h , h = h + 1, and execute step (4);

[0016] (6) Obtain the visual localization detection results:

[0017] Use the test sample set E1 as the input of the trained visual localization network model G* for forward propagation to obtain the visual localization results corresponding to K - M test samples.

[0018] Compared with the prior art, the present invention has the following advantages:

[0019] (1) The localization network model constructed by the present invention includes a text-guided visual feature generation network and a text encoder arranged in parallel. During the training of the model and the process of obtaining the localization results, the global text features are used to guide the generation of visual features at the channel level and the spatial level, making full use of the global semantic information of the text features, reducing the ambiguity in the semantic information, and using text features at different levels to guide the visual features at multiple stages, making full use of the shallow and deep features of the text to supplement the less significant target features in the original feature map of the remote sensing image, effectively improving the accuracy of visual localization of the remote sensing image.

[0020] (2) The localization network model constructed by the present invention includes a text-guided visual feature generation network and a text encoder arranged in parallel. During the training of the model and the process of obtaining the localization results, the text features are used to guide the visual features at different scales in each stage at the spatial level, making better use of the spatial information of the visual feature maps at different scales, and further improving the accuracy of visual localization of the remote sensing image. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 is a flowchart for the implementation of the present invention;

[0022] Figure 2 is a schematic diagram of the overall structure of the network model of the present invention;

[0023] Figure 3 is a schematic diagram of the structure of the text-guided fusion module in the text-guided visual feature generation network of the present invention;

[0024] Figure 4 It is a schematic structural diagram of the multi-modal fusion network of the present invention;

[0025] Figure 5 It is a schematic structural diagram of the language guidance module in the multi-modal fusion network of the present invention;

[0026] Figure 6 It is a schematic structural diagram of the context guidance module in the multi-modal fusion network of the present invention. Detailed implementation manners

[0027] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0028] Refer to Figure 1 , the present invention includes the following steps:

[0029] Step 1) Obtain a training sample set and a test sample set:

[0030] Obtain 19,160 remote sensing images included in the DIOR_RSVG dataset, and annotate the targets and labels in each remote sensing image. Among them, the label includes a bounding box and text. The format of the bounding box is (x, y, w, h), where x and y respectively represent the x and y coordinates of the upper left point of the bounding box, w and h respectively represent the width and height of the bounding box, and the text is a sentence describing the target; form a training sample set R1 with 15,328 remote sensing images and their corresponding labels, and form a test sample set E1 with the remaining 3,832 remote sensing images and their corresponding labels

[0031] Step 2) Construct a remote sensing image visual positioning network model G, and its structure is as Figure 2 shown:

[0032] Construct a remote sensing image visual positioning network model G that includes a text-guided visual feature generation network and a text encoder arranged in parallel, and a multi-modal fusion network and a positioning network cascaded in sequence at the output ends of the text-guided visual feature generation network and the text encoder; the text-guided visual feature generation network includes a downsampling block cascaded in sequence, 4 composite modules composed of cascaded visual feature extraction modules and text-guided fusion modules, and 6 Transformer encoders; the text encoder includes 4 text feature extraction blocks cascaded in sequence, the first text feature extraction block is connected to the corresponding first text-guided fusion module, the second text feature extraction block is connected to the corresponding second text-guided fusion module, the third text feature extraction block is connected to the corresponding third text-guided fusion module, and the fourth text feature extraction block is connected to the corresponding fourth text-guided fusion module; the multi-modal fusion network includes a language-guided module and a context-guided module arranged in parallel, and 6 Transformer decoders cascaded in sequence at the output ends of the language-guided module and the context-guided module;

[0033] The text-guided visual feature generation network, the downsampling block it contains, includes a convolutional layer, a normalization layer, a non-linear activation layer, and a pooling layer stacked in sequence; the first feature extraction block includes 3 residual convolutional blocks connected in sequence, the second feature extraction block includes 4 residual convolutional blocks connected in sequence, the third feature extraction block includes 23 residual convolutional blocks connected in sequence, the fourth feature extraction block includes 3 residual convolutional blocks connected in sequence, and each residual convolutional block includes a 1*1 convolutional layer, a 3*3 convolutional layer, a 1*1 convolutional layer, a normalization layer, and a non-linear activation layer stacked in sequence; the text-guided fusion module includes a channel-level language-guided fusion module and a spatial-level language-guided fusion module connected in sequence, and its structure is as Figure 3 shown: the channel-level language-guided fusion module includes two parallel linear projection layers, and a channel-level multiplication block connected in series with the linear projection layer, and the channel-level multiplication block includes a non-linear activation layer, a convolutional layer, a normalization layer, and a non-linear activation layer connected in sequence; the spatial-level language-guided fusion module includes a linear layer and a non-linear activation layer connected in sequence; the Transformer encoder includes a multi-head self-attention block and a feed-forward network block connected in sequence, the multi-head self-attention block includes a multi-head self-attention layer, a dropout layer, and a normalization layer connected in sequence, and the feed-forward network block includes two linear projection layers, a dropout layer, and a normalization layer connected in sequence;

[0034] The text encoder, the feature extraction block it contains includes 4 Transformer encoders connected in sequence; each Transformer encoder includes a multi-head self-attention block and a feed-forward network block connected in sequence, the multi-head self-attention block contains a multi-head self-attention layer, a dropout layer and a normalization layer connected in sequence, and the feed-forward network block contains two linear projection layers, a dropout layer and a normalization layer connected in sequence;

[0035] The overall structure of the multi-modal fusion network is as Figure 4 shown, and the language guidance module it contains is as Figure 5 shown, including a language guidance processing block and an original feature processing block arranged in parallel, where the language guidance processing block includes a multi-head cross-attention layer, a linear projection layer, and a normalization layer connected in sequence, and the original feature processing block includes a cascaded linear projection layer and a normalization layer; the context guidance module is as Figure 6 shown, including a normalization layer and a context guidance processing block arranged in parallel, where the context guidance processing block includes two multi-head cross-attention modules and a normalization layer cascaded in sequence, and the multi-head cross-attention module contains a multi-head cross-attention layer, a dropout layer and a normalization layer connected in sequence; the Transformer decoder includes a multi-head self-attention block, a multi-head cross-attention block and a feed-forward network block connected in sequence, where the multi-head cross-attention block contains a multi-head cross-attention layer, a dropout layer and a normalization layer connected in sequence;

[0036] The localization network includes two fully connected layers and a non-linear activation layer connected in sequence;

[0037] (3) Initialize parameters:

[0038] Initialize the number of iterations as h, the maximum number of iterations H = 150, and the weight and bias parameters of the visual localization network model G h at the h-th iteration are w h and b h respectively, and let h = 0, G h = G;

[0039] (4) Train the visual localization network model G:

[0040] Randomly select 8 training samples from the training sample set R1 with replacement as the input of the visual localization network model G for forward propagation:

[0041] (4a) The text-guided visual feature generation network extracts features from the pictures in each training sample under text guidance and enhances the obtained visual feature maps to obtain 8 text-guided enhanced visual feature maps; the text encoder extracts features from the text descriptions in each training sample to obtain 8 text features;

[0042] (4a1) In the text-guided visual feature generation network, the downsampling block downsamples each image to obtain 8 downsampled feature maps.

[0043] (4a2) When n = 1, the first visual feature extraction block extracts features from each feature map to obtain 8 visual feature maps; the first text feature extraction block extracts features from the label text corresponding to each feature map to obtain 8 text features; the first text-guided fusion module performs feature fusion at the channel level and spatial level on each visual feature map and the corresponding text feature to obtain 8 text-guided visual feature maps output by the first composite module, and uses them as the input of the visual feature extraction block when n = 2.

[0044] Among them, in the channel-level language-guided fusion block of the text-guided fusion module, two parallel linear projection layers project the input 8 visual feature maps and the corresponding text features into the same dimension as the input of the channel-level multiplication block; the channel-level multiplication block performs an average operation on the sentence feature and word feature of each text feature to obtain 8 global text features P W , and weights the visual feature maps with the global text features at the channel level to obtain 8 channel-level text-guided visual feature maps F v ':

[0045]

[0046] F v ' = σ(F cnn (P w * F v ) + F v )

[0047] Among them, P cls is the [CLS] embedding, AVG() is the average operation function, is N l word features, σ represents the RELU activation function, F v is the visual feature map, F cnn represents a convolutional layer, * represents channel-level multiplication;

[0048] In the spatial-level language-guided fusion block of the text-guided fusion module, first weight and offset the 8 text features to generate 8 groups of dynamic linear layer parameters

[0049]

[0050]

[0051] Among them, P clsdenotes [CLS] embedding, θ() denotes the prediction function, D in is the dimension of the input visual features, D out is the dimension of the output visual features;

[0052] When the feature dimension is high, the number of parameters is extremely large. Therefore, using the idea of matrix factorization, for the parameters are decomposed, and for the parameters are matrix factorized:

[0053]

[0054]

[0055] where U is the dynamic matrix generated according to the text features, W g is the learnable weight parameter, is the transpose matrix of W g b g is the learnable offset parameter, and S is the static learnable matrix;

[0056] For the input 8 channel-level text-guided visual feature maps, use the linear layer generated by the corresponding text features for weighted sum and offset, and then use the Sigmoid function to map it to the attention score to obtain 8 spatial attention scores Score:

[0057]

[0058] where Sigmoid() represents the Sigmoid function;

[0059] Multiply the 8 channel-level text-guided visual feature maps by the corresponding spatial attention scores to refine the visual features at the spatial level and obtain 8 text-guided visual feature maps;

[0060] (4a3) When n = 2, 3, 4, the nth visual feature extraction block extracts features from each text-guided visual feature map output by the (n - 1)th text-guided fusion module to obtain 8 visual feature maps; the nth text feature extraction block extracts features from each text feature output by the (n - 1)th text feature extraction block to obtain 8 text features; the nth text-guided fusion module performs feature fusion at the channel level and spatial level on each visual feature map and the corresponding text feature, a total of 3 times, to obtain 8 text-guided visual feature maps output by the 4th composite module;

[0061] (4a4) 6 Transformer encoders sequentially encode each text-guided visual feature map output by the 4th text-guided fusion module to obtain 8 text-guided enhanced visual feature maps;

[0062] (4b) The multi-modal fusion network performs multi-modal feature fusion on each text-guided enhanced visual feature map output by the text-guided visual feature generation network and the corresponding text features output by the text encoder to obtain 8 multi-modal fusion feature maps;

[0063] (4b1) The language-guided module of the multi-modal fusion network fuses each text-guided enhanced visual feature map and the corresponding text features through multi-head cross-attention to obtain 8 language-guided visual features and projects each text-guided enhanced visual feature map and the corresponding language-guided visual feature to the same dimension, and 8 fine-grained correlations S are obtained through calculation i :

[0064]

[0065]

[0066] where CA() represents the multi-head cross-attention function, F v ” represents the text-guided enhanced feature map, F l represents the text feature; α and σ are fixed coefficients, and exp() is the exponential function with the natural constant e as the base. is the projected text-guided enhanced visual feature map, is the projected language-guided visual feature;

[0067] (4b2) The context-guided module of the multi-modal fusion network first fuses each text-guided enhanced visual feature map and the corresponding text features through multi-head cross-attention to obtain 8 semantic feature-related visual features F y , and then uses the second multi-head cross-attention to fuse each text-guided enhanced visual feature map and the corresponding semantic feature-related visual features to obtain 8 context-guided visual features Finally, pixel-wise multiplication is performed on each context-guided visual feature and the corresponding fine-grained correlation obtained by the language verification module to obtain 8 context-enhanced visual features

[0068] F y = CA(F v ”, F l , F l )

[0069]

[0070]

[0071]

[0072]

[0073] Among them, W Q and W K are learnable weight matrices, and are the transposed matrices of W Q and W K respectively, and Q and K are the Query and Key values of the attention function;

[0074] (4b3) The Transformer decoder in the multimodal fusion module first creates 8 target embeddings t q , where the size of each target embedding is 1*1 and its initial value is 1;

[0075] The input of the i-th Transformer decoder is When i = 1, The i-th Transformer decoder uses multi-head cross-attention for each target embedding and the corresponding text features to obtain 8 text-related target embeddings t' q ; then use multi-head cross-attention for each text-related target embedding and the corresponding text-guided enhanced feature map and context-enhanced visual feature to obtain 8 multimodal fusion target embeddings t' q ', and the feed-forward network fuses each multimodal fusion target embedding with the corresponding original input target embedding to obtain 8 multimodal fusion feature maps This is done 6 times to obtain the 8 multimodal fusion feature maps output by the 6th Transformer decoder :

[0076]

[0077]

[0078]

[0079]

[0080] Among them, t' q is the text-related target embedding, represents the target embedding input to the i-th Transformer decoder, t' q ' represents the multimodal fusion target embedding, LN() is the layer normalization function, and FFN() represents the feed-forward network composed of two linear projection layers, represents the normalized target embedding;

[0081] (4c) The positioning network predicts the coordinates of the annotation boxes in each multi-modal fusion feature map. The format of the annotation box is (x, y, w, h), where x and y represent the x and y coordinates of the upper left point of the annotation box, and w and h represent the width and height of the annotation box, respectively, to obtain 8 visual positioning results, and the positioning results are the coordinates of 8 predicted annotation boxes;

[0082] (5) Update the parameters of the visual positioning network model:

[0083] Using the 8 visual positioning results obtained in step (4), for the visual positioning network model G h Update the weight and bias parameters w h 、b h to obtain the network model G of this iteration h :

[0084] (5a) Using the Smooth-L1 loss function and the GIoU loss function, calculate the model loss value L through the generated 8 visual positioning results, that is, the 8 predicted annotation boxes and the annotation boxes of the corresponding images h :

[0085]

[0086] b = (x, y, w, h)

[0087]

[0088] where b and are the four-dimensional coordinates of the predicted annotation box and the annotation box in the label. x and y represent the x and y coordinates of the upper left point of the coordinate box, and w and h represent the width and height of the coordinate box respectively. L smooth-l1 represents the Smooth-L1 loss, L goiu represents the GIoU loss, and γ is the weight coefficient between the two losses;

[0089] (5b) Calculate the partial derivatives of L h with respect to the weight parameter ω h and the bias parameter b h using the chain rule and Finally, according to update ω h 、b h to obtain the network model G of this iteration h :

[0090]

[0091]

[0092] where ωh , b h represents G h The weights and bias parameters of all learnable parameters, w h ', b h ' represents ω h , b h The updated results of, and lr represents the learning rate.

[0093] Judge whether h > H holds. If so, obtain the trained visual localization network model G*, otherwise, set h = h + 1, G = G h , and execute step (4);

[0094] (6) Obtain the visual localization detection result:

[0095] Use the test sample set E1 as the input of the trained visual localization network model G* for forward propagation to obtain the visual localization results corresponding to 3832 test samples.

[0096] Next, in combination with the simulation experiment, the technical effects of the present invention will be further described:

[0097] 1. Simulation conditions and content:

[0098] The hardware platform for the simulation experiment is: the processor is an Intel(R) Core i9-9900K CPU with a main frequency of 3.5 GHz, the memory is 32 GB, and the graphics card is an NVIDIA GeForce RTX 2080Ti. The software platform for the simulation experiment is: the Ubuntu 16.04 operating system, the python version is 3.7, and the Pytorch version is 1.7.1.

[0099] Conduct a comparative simulation on the positioning accuracy of the present invention and the existing paper RSVG: Exploring Data and Models for Visual Grounding on Remote Sensing Data, and the results are shown in Table 1.

[0100] 2. Analysis of simulation results:

[0101] Referring to Table 1, the evaluation metrics refer to the paper RSVG: Exploring Data and Models for Visual Grounding on Remote Sensing Data, using Pr@0.5, Pr@0.6, Pr@0.7, Pr@0.8, Pr@0.9, where Pr@0.5 means that for a given RS image query pair, if the intersection over union (IoU) with the ground truth bounding box is higher than the threshold of 0.5, the predicted bounding box is considered correct; the specific simulation results are shown in the following table. Compared with the prior art, the visual positioning accuracy of the present invention has been significantly improved:

[0102] 0.5 0.6 0.7 0.8 0.9 Prior art 76.78 72.68 66.74 56.42 35.07 The present invention 80.65 76.92 70.52 59.80 39.07

Claims

1. A text-guided visual positioning method for remote sensing images, characterized in that, It includes the following steps: (1) Obtain a training sample set and a test sample set: Label the targets included in each of the K remote sensing images obtained, and form a training sample set R1 by combining M remote sensing images with their corresponding annotation boxes and texts, and form a test sample set E1 by combining the remaining K - M remote sensing images with their corresponding annotation boxes and texts, where K ≥ 500, (2) Construct a remote sensing image visual localization network model G: Construct a remote sensing image visual localization network model G including a text-guided visual feature generation network and a text encoder arranged in parallel, and a multi-modal fusion network and a localization network cascaded in sequence at the output ends of the text-guided visual feature generation network and the text encoder; the text-guided visual feature generation network includes a downsampling block, N composite modules composed of cascaded visual feature extraction modules and text-guided fusion modules, and R Transformer encoders cascaded in sequence, where the text-guided fusion module includes a channel-level language-guided fusion module and a spatial-level language-guided fusion module connected in sequence; the text encoder includes N text feature extraction blocks cascaded in sequence, and the nth text feature extraction block is also connected to the corresponding nth text-guided fusion module; the multi-modal fusion network includes a language-guided module and a context-guided module arranged in parallel, and S Transformer decoders cascaded in sequence at the output ends of the language-guided module and the context-guided module; where, N≥1, R≥1, S≥1; (3) Initialize the parameters: Initialize the number of iterations as h, the maximum number of iterations as H, where H ≥ 150, and the visual localization network model G at the h-th iteration h has weight and bias parameters w h and b h respectively, and let h = 0, G h = G; (4) Train the visual localization network model G: Randomly select L training samples from the training sample set R1 with replacement as the input of the visual localization network model G for forward propagation to obtain L visual localization results, where, 1≤L≤M; (5) Update the parameters of the visual localization network model: The L visual positioning results obtained through step (4) are used to update the weights and bias parameters w h and b h of the visual positioning network model G h to obtain the network model G h for this iteration. Then, it is judged whether h > H holds. If so, the trained visual positioning network model G* is obtained; otherwise, let G = G h , h = h + 1, and step (4) is executed; (6) Obtain the visual localization detection results: Use the test sample set E1 as the input of the trained visual localization network model G* for forward propagation to obtain the visual localization results corresponding to K-M test samples.

2. The method according to claim 1, wherein In the remote sensing image visual localization network model G described in step (2), where: For the text-guided visual feature generation network, the downsampling block it contains includes a convolutional layer, a normalization layer, a non-linear activation layer, and a pooling layer stacked in sequence; the feature extraction block includes a plurality of residual convolutional blocks connected in sequence, and each residual convolutional block includes a convolutional layer, a normalization layer, and a non-linear activation layer connected in sequence; the channel-level language-guided fusion module includes two parallel linear projection layers and a channel-level multiplication block connected thereto, and the channel-level multiplication block includes a non-linear activation layer, a convolutional layer, a normalization layer, and a non-linear activation layer connected in sequence; the spatial-level language-guided fusion module includes a linear layer and a non-linear activation layer connected in sequence; the Transformer encoder includes a multi-head self-attention block and a feed-forward network block connected in sequence, where the multi-head self-attention block includes a multi-head self-attention layer, a dropout layer, and a normalization layer connected in sequence, and the feed-forward network block includes two linear projection layers, a dropout layer, and a normalization layer connected in sequence; A text encoder, wherein the feature extraction block thereof includes a plurality of Transformer encoders connected in sequence, each of which is composed of a cascaded multi-head self-attention block and a feed-forward network block. The multi-head self-attention block includes a multi-head self-attention layer, a dropout layer, and a normalization layer connected in sequence. The feed-forward network block includes two linear projection layers, a dropout layer, and a normalization layer connected in sequence. A multi-modal fusion network, wherein the language guidance module thereof includes a language guidance processing block and an original feature processing block arranged in parallel. The language guidance processing block includes a multi-head cross-attention layer, a linear projection layer, and a normalization layer connected in sequence. The original feature processing block includes a linear projection layer and a normalization layer connected in cascade. The context guidance module includes a normalization layer and a context guidance processing block arranged in parallel. The context guidance processing block includes two multi-head cross-attention modules and a normalization layer connected in cascade in sequence. A Transformer decoder includes a multi-head self-attention block, a multi-head cross-attention block, and a feed-forward network block connected in sequence. The multi-head cross-attention block includes a multi-head cross-attention layer, a dropout layer, and a normalization layer connected in sequence. A localization network, including two fully connected layers and a non-linear activation layer stacked in sequence.

3. The method according to claim 2, wherein The training of the visual localization network model G in step (4) is implemented as follows: (4a) The text-guided visual feature generation network enhances the features extracted from the image in each training sample under text guidance to obtain L text-guided enhanced visual feature maps. The text encoder extracts features from the text in each training sample to obtain L text features. (4b) The multi-modal fusion network performs multi-modal feature fusion on each text-guided enhanced visual feature map output by the text-guided visual feature generation network and the corresponding text feature output by the text encoder to obtain L multi-modal fusion feature maps. (4c) The localization network predicts the coordinates of the annotation box in each multi-modal fusion feature map to obtain L visual localization results.

4. The method according to claim 3, characterized in that After the text-guided visual feature generation network in step (4a) extracts features from the image in each training sample under text guidance, the implementation steps are as follows: (4a1) The downsampling block in the text-guided visual feature generation network downsamples each image to obtain L downsampled feature maps. (4a2) When n = 1, the first visual feature extraction block extracts features from each feature map to obtain L visual feature maps. The first text feature extraction block extracts features from the label text corresponding to each feature map to obtain L text features. The first text-guided fusion module performs feature fusion at the channel level and the spatial level on each visual feature map and the corresponding text feature to obtain L text-guided visual feature maps output by the first composite module, and uses them as the input of the visual feature extraction block when n = 2. When n = 2...N, the n-th visual feature extraction block extracts features from each text-guided visual feature map output by the (n - 1)-th text-guided fusion module to obtain L visual feature maps; the n-th text feature extraction block extracts features from each text feature output by the (n - 1)-th text feature extraction block to obtain L text features; the n-th text-guided fusion module performs feature fusion at the channel level and spatial level on each visual feature map and the corresponding text feature, a total of N - 1 times, to obtain L text-guided visual feature maps output by the N-th composite module; (4a4) R Transformer encoders encode each text-guided visual feature map output by the N-th text-guided fusion module to obtain L text-guided enhanced visual feature maps.

5. The method according to claim 3, wherein The multi-modal fusion network described in step (4b) performs multi-modal feature fusion on each text-guided enhanced visual feature map output by the text-guided visual feature generation network and the corresponding text feature output by the text encoder. The implementation steps are as follows: (4b1) The language-guided module of the multi-modal fusion network fuses each text-guided enhanced visual feature map and the corresponding text feature to obtain L language-guided visual features, projects each text-guided enhanced visual feature map and the corresponding language-guided visual feature to the same dimension, and calculates L fine-grained correlations; (4b2) The context-guided module of the multi-modal fusion network first fuses each text-guided enhanced visual feature map and the corresponding text feature to obtain L semantic feature-related visual features, then fuses each text-guided enhanced visual feature map and the corresponding semantic feature-related visual feature to obtain L context-guided visual features, and finally performs pixel-wise multiplication on each context-guided visual feature and the corresponding fine-grained correlation obtained by the language verification module to obtain L context-enhanced visual features; (4b3) The Transformer decoder in the multi-modal fusion module creates L target embeddings, and fuses each target embedding with the corresponding text feature, text-guided enhanced feature map, and context-enhanced visual feature to obtain L multi-modal fusion feature maps.

6. The method according to claim 1, wherein The implementation steps for updating the parameters of the visual localization network model described in step (5) are as follows: (5a) Calculate the model loss value L using the Smooth-L1 loss function and the GIoU loss function, based on the L generated visual localization results and the annotation boxes of the corresponding images. h : b = (x, y, w, h) Among them, b and are the four-dimensional coordinates of the labeled box predicted in the visual positioning result and the labeled box in the actual label respectively. x and y represent the x and y coordinates of the upper left point of the coordinate box, and w and h represent the width and height of the coordinate box. L smooth-l1 represents the Smooth-L1 loss, and L goiu represents the GIoU loss. γ is the weight coefficient between the two losses; (5b) Calculate L through the chain rule h For the weight parameter ω h and the bias parameter b h partial derivatives and Finally, according to for ω h , b h perform an update to obtain the network model G for this iteration h : Among them, ω h , b h represent the weights and bias parameters of all learnable parameters of G h , w h ', b h ' represent the updated results of ω h , b h , and lr represents the learning rate.

Citation Information

Patent Citations

  • End-to-end video space-time visual positioning system based on visual language Transform

    CN113849668A

  • Visual positioning method and device based on hierarchical cross-modal context attention mechanism

    CN116152810A