Image-field bidirectional fusion natural language guided unmanned aerial vehicle target retrieval and positioning method, system and medium

By employing a natural language-guided UAV target retrieval and localization method that integrates graph and field fusion, explicitly modeling discrete landmark relationships and implicit environmental layouts, the method addresses the issues of inaccurate target localization and background noise interference in wide-area aerial photography scenarios, achieving higher robustness and accuracy.

CN122636731APending Publication Date: 2026-08-25BEIJING INSTITUTE OF TECHNOLOGY (ZHUHAI)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610899803.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-22
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing vision-language models struggle to adequately respond to small targets, weak targets, and complex orientational relationships in wide-area aerial photography scenarios. Furthermore, drone images exhibit significant variations at different viewpoints and altitudes, leading to inaccurate target localization and substantial background noise interference.

Method used

A natural language-guided UAV target retrieval and localization method is adopted, which integrates graph and field. By constructing a graph-field decoupling and collaborative modeling mechanism, it explicitly models discrete landmark relationships and implicit environmental layout. It also utilizes a bidirectional cross-attention fusion mechanism to optimize cross-modal and multi-view composite problems, achieving collaborative alignment between local target semantics and global scene context.

Benefits of technology

It improves the accuracy of small target localization and the understanding of complex spatial relationships, suppresses background noise interference, and enhances the robustness and interpretability of natural language-guided UAV target retrieval and navigation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122636731A_ABST
    Figure CN122636731A_ABST
Patent Text Reader

Abstract

The application relates to a graph-field bidirectional fusion natural language guided unmanned aerial vehicle target retrieval and positioning method, a system and a medium, the method comprising the following steps: theoretically analyzing a cross-modal and multi-view compound problem in a natural language guided unmanned aerial vehicle target retrieval and positioning task; constructing a graph-field decoupling and collaborative modeling mechanism, establishing a graph-field double-branch parallel modeling framework, specifically comprising constructing a landmark graph branch and an environment field branch; training the graph-field double-branch parallel modeling framework to obtain a graph-field bidirectional fusion strategy model; and training the trained graph-field bidirectional fusion strategy model on a test set to output a target retrieval result and / or a target positioning result under the guidance of natural language. The application can effectively suppress background noise and global semantic deviation in a wide-area aerial scene, improve the accuracy, robustness and interpretability of small target positioning, complex spatial relationship understanding and natural language guided unmanned aerial vehicle target retrieval and navigation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the interdisciplinary fields of UAV intelligent perception, computer vision and natural language processing, and in particular to a method, system and medium for UAV target retrieval and localization guided by natural language through image-field bidirectional fusion. Background Technology

[0002] With the widespread adoption of low-altitude platforms and high-resolution remote sensing imaging equipment, drones are increasingly being used in smart city inspections, disaster response, resource surveys, security monitoring, and geolocation. Compared to traditional methods that rely on manual control or coordinate retrieval, using natural language commands to directly guide drones in target retrieval and localization tasks can significantly improve human-machine interaction efficiency and mission deployment flexibility.

[0003] However, existing vision-language models in wide-area aerial photography scenarios tend to learn globally salient regions of the image, making it difficult to fully respond to small targets, weak targets, and complex orientational relationships described in the text. On the other hand, solutions that rely solely on local region proposals are prone to introducing a large amount of background noise and redundant candidates, thereby weakening the ability to distinguish real targets. Especially in drone scenarios with multiple buildings, roads, and reference points, the text often contains target attributes, reference point relationships, and spatial topological constraints simultaneously, making it difficult for existing models to collaboratively model the "target object ontology" and the "layout of the target's environment."

[0004] On the other hand, UAV images from different viewpoints and altitudes exhibit significant scale variations, imaging distortion, occlusion, and visibility differences. The representation of the same target is prone to drift under different observation conditions, making it difficult for retrieval methods based on single vector similarity to operate stably. Therefore, there is an urgent need for a technical solution that can simultaneously explicitly model discrete landmark relationships, implicitly model continuous environmental fields, and utilize multi-view evidence to verify and correct natural language semantic hypotheses. Summary of the Invention

[0005] This invention provides a graph-field bidirectional fusion natural language-guided UAV target retrieval and localization method, system, and medium. It aims to simultaneously and explicitly model discrete landmark relationships and implicit environmental layout information, achieving coordinated alignment between local target semantics and global scene context. This effectively suppresses background noise and global semantic bias in wide-area aerial photography scenarios, improving the accuracy, robustness, and interpretability of small target localization, complex spatial relationship understanding, and natural language-guided UAV target retrieval and navigation. It solves the problems of global semantic bias, weak small target localization capability, large background noise interference, and difficulty in modeling complex spatial relationships in existing technologies in wide-area aerial photography scenarios.

[0006] This invention provides a natural language-guided target retrieval and localization method for unmanned aerial vehicles (UAVs) that integrates graph and field fusion. The method includes the following steps: Step S10: A theoretical analysis is conducted on the cross-modal and multi-perspective complex challenges in natural language-guided UAV target retrieval and localization tasks. Step 20: Construct a graph-field decoupling and collaborative modeling mechanism to prove that combining discrete landmark relationship modeling with global environment layout modeling can optimize the cross-modal, multi-view composite problem; Step S30: Establish a graph-field dual-branch parallel modeling framework for natural language guided UAV missions, specifically including: constructing a landmark map branch using a map salient target display modeling method, and constructing an environmental field branch using an environmental macro-layout implicit modeling method; Step 40: Train the graph-field dual-branch parallel modeling framework, and use a bidirectional cross-attention fusion mechanism to jointly optimize the landmark graph branch and the environmental field branch to obtain a graph-field bidirectional fusion strategy model for natural language-guided UAV target retrieval and localization. Step S50: Train the trained graph-field bidirectional fusion strategy model on the test set and output the target retrieval results and / or target localization results guided by natural language.

[0007] A further technical solution of the present invention is that step S10 includes: Step S101: Define the natural language-guided UAV target retrieval and localization task as a cross-modal semantic alignment task based on UAV viewpoint images, with the input being the UAV viewpoint image. With natural language instructions Among them, the drone perspective image satisfies Natural language instructions It includes target entities, attribute descriptions, and spatial relationship semantics; the goal of the localization task is to locate objects in a visual dataset. With text datasets The above establishes a bidirectional cross-modal correspondence between images and text to achieve text-guided target retrieval, image-guided text retrieval, and target localization. Step S102 defines the natural language-guided UAV target retrieval and localization task as a unified retrieval and localization problem. The retrieval subtask retrieves the most relevant image from a candidate UAV image set based on natural language instructions, or queries the most relevant text based on an image. The localization subtask determines the location of the target region in the image based on natural language instructions. The retrieval subtask learns a graph-text similarity function in a joint embedding space. The result of the search is expressed as follows: , in, Represents a visual dataset. Represents a text dataset. Represents a visual encoder. Indicates a text encoder; Step S103, the localization subtask is defined as: for natural language instructions The target region is pointed to, and its prediction in the image is made. bounding box in ,in, Indicates the center coordinates of the target area. and These represent the width and height of the bounding box, respectively; the bounding box is predicted by solving the equation. With the true bounding box The deviation between them is used to achieve natural language-guided target localization, and the localization loss is expressed as: ; Step S104 further defines the natural language-guided UAV task as a cross-modal, multi-view coupling problem. Cross-modal coupling means that the local target semantics in the natural language description needs to establish an accurate correspondence with the local area in the wide-area aerial image. Multi-view coupling means that the same target has significant representational shifts under different imaging angles, scales, and background conditions. Therefore, the core of the localization subtask is to establish an accurate correspondence between discrete landmarks and combined semantics in a large-scale environmental background. Step S105: Based on the definition of the localization subtask, it is clarified that the method needs to simultaneously meet two types of modeling requirements: local salient target recognition and global environment layout perception. Local salient target recognition is used to enhance the sensitivity to discrete targets indicated by the text and their spatial relationships, while global environment layout perception is used to provide the macroscopic context of the scene where the target is located, thereby providing a problem definition basis for subsequent graph-field dual-branch parallel modeling.

[0008] A further technical solution of the present invention is that step S20 includes: Step S201 addresses the cross-modal granularity mismatch problem in natural language-guided UAV target retrieval and localization tasks, where "text only covers local target semantics, while images contain a large range of complex environmental backgrounds." A graph-field decoupling and collaborative modeling mechanism is constructed, decomposing visual representation into two parallel pathways: discrete landmark relationship modeling and global environment layout modeling. Discrete landmark relationship modeling captures the local targets referred to in the text and their spatial topological relationships, while global environment layout modeling captures the macro-context formed by the distribution of roads, terrain, building clusters, and repetitive textures. Through bidirectional interactive fusion, the two types of representations form complementary constraints within a unified embedding space, thereby suppressing background interference and enhancing the semantic saliency of small targets. Step S202 demonstrates how discrete landmark relationship modeling improves model performance, specifically including: Step S021: Establish a local alignment propagation theory for the discrete landmark relationship modeling branch, defining the visual graph and text graph in the first step. The node embeddings of the layers are respectively and The final fused visual representation is defined as follows: , in, This represents the output features of the visual backbone network. Indicate the number of propagation layers in the graph neural network; and assume the message passing function. Fusion function with cross attention It satisfies the Lipschitz continuity condition, that is: , , It is also assumed that the graph topological differences are constrained by an upper bound on the semantic differences of the nodes, that is: , in, This represents the adjacency structure of the visual graph. This represents the adjacency structure of a text graph. Indicates the structural coupling coefficient; Step S2022, under the assumptions of step S2021, proves that the local node-level alignment error can propagate along the inter-layer paths of the graph neural network and ultimately constrain the global cross-modal representation differences; wherein, the initial node-level alignment error is defined as: , The final cross-modal representation differences satisfy: , in, Represents visual backbone feature noise. This indicates irreducible modal differences; thus, it shows that when node-level semantic alignment error... During descent, the final fusion representation and The cross-modal deviation is reduced synchronously, which theoretically proves that discrete landmark relationship modeling can propagate local semantic alignment into global representation consistency. Step S2023, further prove the inter-layer error propagation in step S2022, and define the first... The differences in layer graph node embedding are: , Then, from the decomposition of the characteristic error term and the structural error term, we can obtain: , Combining the structural consistency assumption, we can further obtain: , We can obtain the following through recursion: , Substituting the Lipschitz upper bound of the cross-attention fusion function, we obtain the final cross-modal representation difference upper bound, thus completing the theoretical proof of the discrete landmark relationship modeling branch; Step S203 demonstrates how global environment layout modeling improves model performance, specifically including: Step S2031: Establish frequency domain alignment propagation theory for the global environment layout modeling branch, and define the visual backbone output as... Its frequency domain representation is defined as: ,in, Represent the Discrete Fourier Transform; let the adaptive frequency mixing operator be... The modulated visual spectrum is then: And the environmental field representation is reconstructed through inverse transformation: Let the ideal environment field, implicitly defined by the text global vector, be represented as... The corresponding ideal spectrum is: The theoretical goal of the environmental field branch is to make and Alignment; Step S2032: Establish the frequency domain and spatial domain error equivalence relationship for step S2031, assuming that the discrete Fourier transform satisfies Parseval's theorem, that is, for any signal... and ,have: , And it is assumed that the frequency mixing operator satisfies Lipschitz continuity: , Meanwhile, assume that the spectrum alignment residual satisfies: , in, Indicates frequency domain alignment error; Step S2033, under the assumptions of step S2032, prove that frequency domain alignment can propagate into spatial domain environmental field uniformity; let The spatial inconsistency of the global environment field satisfies: , in, This represents the approximation error of the ideal environmental field; further considering the visual backbone input noise, we have: , in, This represents the ideal visual input; thus proving that the frequency domain alignment error... The optimization will linearly improve the consistency of the spatial domain environment field, theoretically demonstrating that global environment layout modeling can effectively enhance the robustness of cross-modal alignment; Step S204: Based on the theoretical results of steps S202 and S203, the inference of the graph-field decoupling and collaborative modeling mechanism is obtained: discrete landmark relationship modeling is responsible for propagating node-level semantic alignment error constraints to the global fusion representation, improving target-level and relationship-level identification capabilities; global environment layout modeling is responsible for propagating spectral alignment errors into spatial field consistency, improving scene-level and context-level robustness.

[0009] A further technical solution of the present invention is that step S30 includes: Step S301: Construct a graph-field dual-branch parallel modeling framework. Perform local salient target modeling and global environment layout modeling on the input UAV view image and natural language commands, respectively. The local salient target modeling forms a landmark map branch; the global environment layout modeling forms an environment field branch. The landmark map branch is used to explicitly represent target entities and their spatial topological relationships, and the environment field branch is used to implicitly represent scene texture distribution, repeating structures and macroscopic layout information. The two branches are processed in parallel and output landmark map features and environment field features respectively. The specific steps for constructing the landmark map branch include: Step S3021: Extract the target from the UAV image using a text-guided target detector in the visual domain. Identify candidate salient landmark areas and construct a visual graph using the bounding boxes of each candidate area as nodes. in, Represents the set of nodes in a visual graph. Represents the set of edges in a visual graph; to characterize the relative spatial relationships between landmarks, nodes are... With nodes The relative center position and scale information between them are encoded into the edge attributes, and its spatial edge features are defined as follows: , in, and Representing nodes respectively With nodes The center coordinates of the corresponding bounding box These represent the width and height of the corresponding bounding box, respectively; through the above spatial edge features, explicit modeling of direction, distance, and scale changes in the visual landmark map is achieved.

[0010] Step S3022: In the landmark graph branch, graph inference based on edge attribute enhancement is performed on the visual graph nodes. The TransformerConv operator is used to aggregate the features of neighboring nodes and edge space features, so that each visual node obtains the topological context related to its neighboring landmarks. The update rule is as follows: , in, Indicates the first Layer nodes Visual features Represents a node The neighborhood set, Represents a node For nodes Attention weights , and The learnable parameter matrix is ​​represented; the update rule enables visual graph nodes to retain explicit spatial geometric constraints while aggregating neighborhood information. Step S3023: Construct a text graph corresponding to the visual graph in the text field, denoted as... ,in, Represents a set of text graph nodes. This represents the text graph edge set. A large language model is used to extract landmark entities, attribute modifiers, and relational descriptions from natural language instructions to initialize text node features. Directional encoding is also used to initialize text graph edge features to characterize the syntactic dependencies and spatial relationships between different semantic entities. Furthermore, a graph attention network is employed to perform semantic propagation on the text graph, with attention coefficients and node update rules as follows: , , , in, Indicates the first Layer nodes Textual features, This represents vector concatenation. and Indicates learnable parameters, It represents a non-linear activation function; through the text graph construction and update process, phrase-level semantics are aggregated into a contextualized landmark semantic representation containing target, attribute, and relation information; Step S3024: Align and represent the visual node features after visual graph inference with the text node features after text graph propagation to obtain the landmark map feature token output by the landmark map branch. This token is used to explicitly characterize the discrete targets corresponding to the natural language instructions, the spatial relationships between targets, and the local geometric topology. The landmark map features are used to enhance the model's ability to recognize spatial relationship descriptions such as center, left, right, top, bottom, and adjacency, and to improve the distinguishability between small targets and non-dominant targets. The specific steps for constructing the environmental field branch include: Step S3031: Process the global feature map output by the visual backbone network. Perform serialization processing on the input feature map Flattened to a length of sequence representation It is then mapped to the frequency domain using a one-dimensional real discrete Fourier transform, and its frequency domain representation is as follows: , in, Represents the imaginary unit. It represents the frequency index; by transforming a two-dimensional spatial layout into a one-dimensional frequency domain sequence, it enables a global perception of wide-area environmental characteristics.

[0011] Step S3032, in the environmental field branch, to adaptively enhance the frequency band related to the environmental layout, generate... A learnable one-dimensional frequency band mask is used, and weighted modulation is performed on each frequency band; where, the first... A frequency band mask is defined as: , in, Indicates the first The center frequency of each frequency band Indicates the first The bandwidth parameters of each frequency band are determined; further, the frequency domain features are modulated element-wise with the frequency band mask and the learnable complex weight matrix to obtain the enhanced frequency domain features: , in, Indicates the first The complex weight matrix corresponding to each frequency band This represents the Hadamard element-wise product; through multi-band spectral enhancement, the model can strengthen macro-environmental patterns such as road grids, building cluster orientations, terrain trends, and repetitive textures. Step S3033: Perform inverse Fourier transform on the enhanced frequency domain features to reconstruct the frequency domain enhancement result back to the spatial domain, and obtain the environmental field branch output features through residual connection and projection mapping. The reconstruction process is expressed as follows: , in, Indicates the inverse Fourier transform. This indicates a tensor shape reconstruction operation. This represents a multilayer perceptron mapping; in some implementations, the environmental field branch further refines the reconstructed spatial features by combining window attention or shifted window attention, in order to take into account both global layout perception and local spatial details. Step S3034: Define the output of the environmental field branch as an environmental field feature token, which is used to implicitly represent the continuous environmental field semantics in the UAV image, including the distribution of roads, terrain, green space, building clusters, and scene layout information formed by repeated textures; the environmental field features and the landmark map features complement each other at the semantic level, wherein the former provides global scene context and the latter provides discrete landmark anchors directly related to the text, thus jointly constituting a graph-field dual-branch parallel modeling framework.

[0012] A further technical solution of the present invention is that step S40 includes: Step S401: Perform bidirectional cross-attention fusion on the landmark map features output by the landmark map branch and the environmental field features output by the environmental field branch to construct a geometry-guided bidirectional cross-attention fusion module; wherein, the bidirectional cross-attention fusion module is used to introduce the global layout context represented by the environmental field branch while maintaining the fine-grained discrete target recognition capability of the landmark map branch, and realize landmark saliency enhancement and environmental region focusing through bidirectional information interaction, thereby forming a unified cross-modal fusion representation suitable for natural language-guided UAV target retrieval and localization tasks; The specific steps for constructing the bidirectional cross-attention fusion module include: Step S4011: Set the output of the landmark map branch to a landmark token sequence. The environment field branch output is an environment token sequence. Considering the high spatial resolution of the environment field branch, to reduce the computational complexity of cross-attention, a representative subset is first uniformly sampled from the environment token sequence. ,in, This represents a subset of environment tokens that participate in the cross-attention calculation; Step S4012: To maintain spatial consistency during graph-field fusion, a geometric bias term is introduced in the calculation of cross-attention weights. This allows the model to prioritize context regions that are physically closer to the current target; where, the first The first landmark node and the first The geometric offset between environment tokens is defined as follows: , in, Indicates the first Normalized spatial coordinates of each landmark node Indicates the first The normalized spatial coordinates of an environment token. This represents the learnable distance attenuation coefficient; the geometric bias term is used to penalize long-distance interactions, making the fusion process more consistent with the local spatial consistency constraints in the UAV perspective scene. Step S4013: In the bidirectional cross-attention fusion module, the context injection process for the landmark query environment is first executed, i.e., Landmark-to-Environment or... In this stage, the landmark token output by the landmark map branch serves as the query, and the environment token output by the environment field branch serves as the key and value. This allows discrete landmarks to absorb relevant global contextual information, thereby enhancing target discrimination capabilities. Its update form is expressed as: , in, This represents the query matrix obtained by linear mapping of landmark tokens. and These represent the key matrix and value matrix obtained by linear mapping from the environment token, respectively. The channel dimension representing the attention head; Step S4014: In the bidirectional cross-attention fusion module, the target location process of the environmental query landmark is further executed. In this process, the environmental token output by the environmental field branch is used as the Query, and the landmark token output by the landmark map branch is used as the Key and Value. This enables the environmental field representation to reverse focus and reweight the significant target area based on the distribution of key landmarks. Its update form is expressed as follows: , in, This represents the query matrix obtained by linear mapping of the environment token. and These represent the key matrix and value matrix obtained by linear mapping of landmark tokens, respectively; through the reverse aggregation process, the environmental field features are further transformed from a global background description into a contextualized background representation constrained by key targets; Step S4015: Update the landmark features obtained in step S4013. The updated environment features obtained in step S4014 By performing stitching, projection mapping, or multilayer perceptron transformation, the image-field fusion feature representation is obtained. The graph-field fusion feature contains both the local structural information of discrete targets and the global layout information of continuous environmental fields, which are used for subsequent graph-text matching, target localization, and multi-view consistent reasoning. The bidirectional interaction mechanism differs from the unidirectional injection fusion strategy and can maintain both the fineness of point target representation and the stability of field layout representation. Step S402: Based on the graph-field fusion feature representation, jointly train the entire dual-branch parallel modeling framework to construct the total loss function. This enables the model to simultaneously meet the requirements of global graph-text semantic alignment, target localization accuracy, and graph structure semantic consistency; the total loss function is expressed as: , in, This indicates the learning loss due to the comparison between text and images. This represents the image-text matching loss. This represents the target bounding box localization loss. This represents the graph structure alignment loss. and This represents the weighting coefficient of the corresponding loss term; The target bounding box localization loss The sum of the bounding box regression loss and the generalized intersection-union (OCU) loss is used to constrain the localization accuracy of the landmark map branch for the text-indicated target. Its expression is: , in, Indicates the predicted bounding box. Represents the true bounding box; The graph structure alignment loss The expression used to constrain the consistency between visual and text graphs in terms of node semantic distribution, spatial relationships, node discriminativeness, and structural dependencies is: , in, , , and These are the weighting coefficients for the corresponding loss terms; the cross-modal distribution alignment loss The expression used to narrow the difference between text node embedding distribution and visual node embedding distribution is: , in, and These represent the probability distributions of text node embedding and visual node embedding, respectively. Spatial Relationship Regression Loss The expression used to constrain the consistency between the predicted spatial relationship vectors between nodes in the visual graph and their true relative coordinates is: , in, Represents the predicted spatial relationship vector. Represents true relative coordinates, Indicates the number of edges in the visual graph; Node-level comparison loss To enhance the discriminative power of semantically matching node pairs, making positively matching node pairs closer in the embedding space, its expression is: , in, This represents the set of positively matching node pairs. Indicates the temperature coefficient. Represents the similarity function; Structural consistency loss The expression used to constrain the consistency between the text graph dependency matrix and the visual graph adjacency matrix is: , in, Represents the text graph dependency matrix. Represents the adjacency matrix of the visual graph. Represents the Frobenius norm; through the stated This enables the landmark map branches to have stronger cross-modal structural alignment capability before fusion; Step S403: The visual encoder, text encoder, landmark map branch, environmental field branch, and bidirectional cross-attention fusion module are jointly optimized end-to-end through the total loss function to obtain a graph-field bidirectional fusion strategy model for natural language-guided UAV target retrieval and localization tasks.

[0013] A further technical solution of the present invention is that step S50 includes: inputting the test set into the pre-trained model, and the model outputting the predicted answer.

[0014] To achieve the above objectives, the present invention also proposes a graph-field bidirectional fusion natural language-guided UAV target retrieval and localization system. The system includes a memory, a processor, and a graph-field bidirectional fusion natural language-guided UAV target retrieval and localization program stored on the processor. The graph-field bidirectional fusion natural language-guided UAV target retrieval and localization program is executed by the processor to perform the steps of the method described above.

[0015] To achieve the above objectives, the present invention also proposes a computer-readable storage medium storing a graph-field bidirectional fusion natural language-guided UAV target retrieval and localization program, wherein the graph-field bidirectional fusion natural language-guided UAV target retrieval and localization program is executed by a processor to perform the steps of the method described above.

[0016] The beneficial effects of the image-field bidirectional fusion natural language-guided UAV target retrieval and localization method, system, and medium of this invention are: (1) This invention uses a dual-path collaborative modeling mechanism of “landmark map + environment field” to explicitly depict the topological relationship between discrete targets and implicitly depict the layout information of continuous scene. Compared with the scheme of using only global features or only local areas, it can more accurately respond to the targets, reference objects and spatial relationships involved in natural language.

[0017] (2) The present invention adopts a bidirectional cross-attention fusion mechanism with geometric bias, which enables local landmarks to actively index the global background context, and enables the global environment field to focus on key landmarks in reverse, thereby effectively suppressing background noise in wide-area aerial images and enhancing the ability to distinguish small targets.

[0018] (3) The present invention adopts a joint optimization method that combines cross-modal comparison, cross-modal matching, localization supervision and graph structure alignment, which is conducive to improving semantic alignment accuracy, spatial reasoning ability and target localization accuracy at the same time. It is applicable to various application scenarios such as natural language guided UAV retrieval, UAV navigation, target matching and geolocation. Attached Figure Description

[0019] Figure 1 This is a flowchart illustrating a preferred embodiment of the natural language-guided UAV target retrieval and localization method based on graph-field bidirectional fusion of the present invention. Figure 2 This is a model framework diagram of a natural language-guided UAV target retrieval and localization method based on graph-field bidirectional fusion; Figure 3 This is a schematic diagram of the environmental field branch modeling principle; Figure 4 This is a schematic diagram of the bidirectional cross-attention fusion module; Figure 5 This is a diagram illustrating the reasoning effect of a natural language-guided drone target retrieval and localization problem.

[0020] Figure 6 This is a schematic diagram of the hardware architecture of the natural language-guided UAV target retrieval and positioning system based on graph-field bidirectional fusion, as described in this invention. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0022] This invention proposes a graph-field dual fusion method for target retrieval and localization of Natural Language-Guided Drones (NLGD). The technical solution adopted in this invention is geared towards NLGD scenarios and constructs a graph-field dual fusion strategy (GFDFFS), including: Step 1, acquiring natural language commands and query images and / or multi-view image sequences collected by the drone; Step 2, performing hierarchical semantic parsing on the natural language commands, extracting target entities, attribute modifiers, directional relationships, and constraint phrases to construct a text semantic graph; Step 3, extracting candidate landmark regions based on text-guided target detection, constructing a visual landmark map containing relative position, direction, and scale relationships, and obtaining landmark map features through a graph neural network (GNN) to form a landmark graph branch (LGB); Step 4, performing serialized frequency domain transformation, multi-band adaptive spectral enhancement under Discrete Fourier Transform (DFT), and field reconstruction on the global visual features to obtain environmental field features, forming an environment field branch (Environment Field Branch). Step 5: Using a bidirectional cross-attention fusion module with geometric bias (CAF), the landmark map features and environmental field features are bidirectionally interacted and collaboratively enhanced to obtain the map-field fusion features; Step 6: Based on the map-field fusion features, image-text matching scoring, target localization, and multi-view evidence consistency correction are performed to output the target retrieval results or localization results.

[0023] Specifically, such as Figures 1 to 5 As shown, a preferred embodiment of the graph-field bidirectional fusion natural language-guided UAV target retrieval and localization method of the present invention includes the following steps: Step S10 involves a theoretical analysis of the complex challenges of cross-modal and multi-perspective tasks in natural language-guided UAV target retrieval and localization.

[0024] Step 20: Construct a graph-field decoupling and collaborative modeling mechanism from a theoretical perspective, and prove that combining discrete landmark relationship modeling with global environment layout modeling can optimize the cross-modal, multi-view composite problem.

[0025] Step S30: Establish a graph-field dual-branch parallel modeling framework for Natural Language Guided Drone (NLGD) missions. Specifically, this includes: constructing a landmark graph branch using a map salient target display modeling method, and constructing an environmental field branch using an environmental macro-layout implicit modeling method.

[0026] Step 40: Train the graph-field bi-branch parallel modeling framework, and use a bi-directional cross-attention fusion mechanism to jointly optimize the landmark graph branch and the environmental field branch to obtain a graph-field bi-directional fusion strategy model for natural language-guided UAV target retrieval and localization.

[0027] Step S50: Train the trained graph-field bidirectional fusion strategy model on the test set and output the target retrieval results and / or target localization results guided by natural language.

[0028] Specifically, step S10 includes: Step S101: Define the natural language-guided UAV target retrieval and localization task as a cross-modal semantic alignment task based on UAV viewpoint images, with the input being the UAV viewpoint image. With natural language instructions Among them, the drone perspective image satisfies .

[0029] Natural Language Commands It includes target entities, attribute descriptions, and spatial relationship semantics. The goal of this task is to: [include] visual datasets. With text datasets The system establishes a bidirectional cross-modal correspondence between images and text to enable text-guided target retrieval, image-guided text retrieval, and target localization.

[0030] Step S102 defines the task as a unified retrieval and localization problem. The retrieval subtask retrieves the most relevant image from a candidate UAV image set based on natural language instructions, or queries the most relevant text based on an image. The localization subtask determines the location of a target region in an image based on natural language instructions. The retrieval subtask learns a graph-text similarity function in a joint embedding space. The optimal retrieval result is expressed as follows: , in, Represents a visual dataset. Represents a text dataset, Represents a visual encoder. This indicates a text encoder.

[0031] Step S103, the localization subtask is defined as: for natural language instructions The target region is pointed to, and its prediction in the image is made. bounding box in in, Indicates the center coordinates of the target area. and These represent the width and height of the bounding box, respectively; the bounding box is predicted by solving the equation. With the true bounding box The deviation between them is used to achieve target localization guided by natural language. The localization loss can be expressed as: .

[0032] Step S104 further defines the natural language-guided UAV task as a cross-modal, multi-view coupling problem. Cross-modal coupling means that the local target semantics in the natural language description needs to establish an accurate correspondence with the local area in the wide-area aerial image. Multi-view coupling means that the same target has significant representational shifts under different imaging angles, scales and background conditions. Therefore, the core of the task is to establish an accurate correspondence between discrete landmarks and combined semantics in a large-scale environmental background.

[0033] Step S105: Based on the task definition, it is clarified that the method needs to simultaneously meet two types of modeling requirements: local salient target recognition and global environment layout perception. Local salient target recognition is used to enhance the sensitivity to discrete targets referred to in the text and their spatial relationships, while global environment layout perception is used to provide the macroscopic context of the scene where the target is located, thereby providing a problem definition basis for subsequent graph-field dual-branch parallel modeling.

[0034] Step S20 specifically includes: Step S201 addresses the cross-modal granularity mismatch problem in natural language-guided UAV target retrieval and localization tasks, where "text only covers local target semantics, while images contain a large range of complex environmental backgrounds." A graph-field decoupling and collaborative modeling mechanism is constructed, decomposing visual representation into two parallel pathways: discrete landmark relationship modeling and global environment layout modeling. Discrete landmark relationship modeling captures the local targets referred to in the text and their spatial topological relationships, while global environment layout modeling captures the macro-context formed by the distribution of roads, terrain, building clusters, and repetitive textures. Through bidirectional interactive fusion, the two types of representations form complementary constraints within a unified embedding space, thereby suppressing background interference and enhancing the semantic saliency of small targets.

[0035] Step S202 demonstrates how discrete landmark relationship modeling improves model performance, specifically including: Step S021: Establish a local alignment propagation theory for the discrete landmark relationship modeling branch, defining the visual graph and text graph in the first step. The node embeddings of the layers are respectively and The final fused visual representation is defined as follows: , in, This represents the output features of the visual backbone network. Indicate the number of propagation layers in the graph neural network; and assume the message passing function. Fusion function with cross attention It satisfies the Lipschitz continuity condition, that is: , , It is also assumed that the graph topological differences are constrained by an upper bound on the semantic differences of the nodes, that is: , in, This represents the adjacency structure of the visual graph. This represents the adjacency structure of a text graph. This represents the structural coupling coefficient.

[0036] Step S2022, under the assumptions of step S2021, proves that the local node-level alignment error can propagate along the inter-layer paths of the graph neural network and ultimately constrain the global cross-modal representation differences; wherein, the initial node-level alignment error is defined as: , The final cross-modal representation differences satisfy: , in, Represents visual backbone feature noise. This indicates irreducible modal differences; thus, it shows that when node-level semantic alignment error... During descent, the final fusion representation and The cross-modal deviation is reduced synchronously, which theoretically proves that discrete landmark relationship modeling can propagate local semantic alignment into global representation consistency.

[0037] Step S2023, further prove the inter-layer error propagation in step S2022, and define the first... The differences in layer graph node embedding are: , Then, from the decomposition of the characteristic error term and the structural error term, we can obtain: , Combining the structural consistency assumption, we can further obtain: , We can obtain the following through recursion: , Substituting the Lipschitz upper bound of the cross-attention fusion function, we obtain the final cross-modal representation difference upper bound, thus completing the theoretical proof of the discrete landmark relationship modeling branch.

[0038] Step S203 demonstrates how global environment layout modeling improves model performance, specifically including: Step S2031: Establish frequency domain alignment propagation theory for the global environment layout modeling branch, and define the visual backbone output as... Its frequency domain representation is defined as: ,in, Represent the Discrete Fourier Transform; let the adaptive frequency mixing operator be... The modulated visual spectrum is then: And the environmental field representation is reconstructed through inverse transformation: Let the ideal environment field, implicitly defined by the text global vector, be represented as... The corresponding ideal spectrum is: The theoretical goal of the environmental field branch is to make and Alignment.

[0039] Step S2032: Establish the frequency domain and spatial domain error equivalence relationship for step S2031, assuming that the discrete Fourier transform satisfies Parseval's theorem, that is, for any signal... and ,have: , And it is assumed that the frequency mixing operator satisfies Lipschitz continuity: , Meanwhile, assume that the spectrum alignment residual satisfies: , in, This indicates the frequency domain alignment error.

[0040] Step S2033, under the assumptions of step S2032, prove that frequency domain alignment can propagate into spatial domain environmental field uniformity; let The spatial inconsistency of the global environment field satisfies: , in, This represents the approximation error of the ideal environmental field; further considering the visual backbone input noise, we have: , in, This represents the ideal visual input; thus proving that the frequency domain alignment error... The optimization will linearly improve the consistency of the spatial domain environment field, theoretically demonstrating that global environment layout modeling can effectively enhance the robustness of cross-modal alignment.

[0041] Step S204: Based on the theoretical results of steps S202 and S203, the inference of the graph-field decoupling and collaborative modeling mechanism is obtained: discrete landmark relationship modeling is responsible for propagating node-level semantic alignment error constraints to the global fusion representation, improving target-level and relationship-level identification capabilities; global environment layout modeling is responsible for propagating spectral alignment errors to spatial field consistency, improving scene-level and context-level robustness; the combination of the two can simultaneously optimize local target saliency and global scene consistency, thus outperforming the scheme of using only a single local modeling or a single global modeling.

[0042] Step S30 specifically includes: Step S301: Construct a graph-field dual-branch parallel modeling framework. Local salient target modeling and global environment layout modeling are performed on the input UAV view image and natural language commands, respectively. Local salient target modeling forms a Landmark Graph Branch (LGB); global environment layout modeling forms an Environment Field Branch (EFB). The Landmark Graph Branch is used to explicitly represent target entities and their spatial topological relationships, while the Environment Field Branch is used to implicitly represent scene texture distribution, repeating structures, and macroscopic layout information. The two branches are processed in parallel and output Landmark Graph features and Environment Field features respectively. The proposed graph-field dual-branch parallel modeling framework is as follows: Figure 2 As shown.

[0043] The specific steps for constructing the landmark map branch include: Step S3021: Extract the target from the UAV image using a text-guided target detector in the visual domain. Identify candidate salient landmark regions and construct a visual graph using the bounding boxes of each candidate region as nodes. in, Represents the set of nodes in a visual graph. Represents the set of edges in a visual graph; to characterize the relative spatial relationships between landmarks, nodes are... With nodes The relative center position and scale information between them are encoded into the edge attributes, and its spatial edge features are defined as follows: , in, and Representing nodes respectively With nodes The center coordinates of the corresponding bounding box These represent the width and height of the corresponding bounding box, respectively; through the above spatial edge features, explicit modeling of direction, distance, and scale changes in the visual landmark map is achieved.

[0044] Step S3022: In the landmark graph branch, graph inference based on edge attribute enhancement is performed on the visual graph nodes. The TransformerConv operator is used to aggregate the features of neighboring nodes and edge space features, so that each visual node obtains the topological context related to its neighboring landmarks. The update rule is as follows: , in, Indicates the first Layer nodes Visual features Represents a node The neighborhood set, Represents a node For nodes Attention weights , and The learnable parameter matrix is ​​represented; the update rule enables visual graph nodes to retain explicit spatial geometric constraints while aggregating neighborhood information.

[0045] Step S3023: Construct a text graph corresponding to the visual graph in the text field, denoted as... ,in, Represents a set of text graph nodes. This represents the text graph edge set. A large language model is used to extract landmark entities, attribute modifiers, and relational descriptions from natural language instructions to initialize text node features. Directional encoding is also used to initialize text graph edge features to characterize the syntactic dependencies and spatial relationships between different semantic entities. Furthermore, a graph attention network is employed to perform semantic propagation on the text graph, with attention coefficients and node update rules as follows: , , , in, Indicates the first Layer nodes Textual features, This represents vector concatenation. and Indicates learnable parameters, The nonlinear activation function is represented; through the text graph construction and update process, phrase-level semantics are aggregated into contextualized landmark semantic representations containing target, attribute, and relation information.

[0046] Step S3024 involves aligning and representing the visual node features obtained through visual graph inference with the text node features obtained through text graph propagation to obtain landmark graph feature tokens output by the landmark graph branch. These tokens are used to explicitly characterize the discrete targets corresponding to natural language instructions, the spatial relationships between targets, and the local geometric topology. The landmark graph features enhance the model's ability to recognize spatial relationships such as center, left, right, top, bottom, and adjacency, and improve the distinguishability between small targets and non-dominant targets. This type of output is represented as Landmark Graph Tokens and used for subsequent cross-branch fusion.

[0047] The specific steps for constructing the environmental field branch include: Step S3031: Process the global feature map output by the visual backbone network. Perform serialization processing on the input feature map Flattened to a length of sequence representation It is then mapped to the frequency domain using a one-dimensional real discrete Fourier transform, and its frequency domain representation is as follows: , in, Represents the imaginary unit. It represents the frequency index; by transforming a two-dimensional spatial layout into a one-dimensional frequency domain sequence, it enables a global perception of wide-area environmental characteristics.

[0048] Step S3032, in the environmental field branch, to adaptively enhance the frequency band related to the environmental layout, generate... A learnable one-dimensional frequency band mask is used, and weighted modulation is performed on each frequency band; where, the first... A frequency band mask is defined as: , in, Indicates the first The center frequency of each frequency band Indicates the first The bandwidth parameters of each frequency band are determined; further, the frequency domain features are modulated element-wise with the frequency band mask and the learnable complex weight matrix to obtain the enhanced frequency domain features: , in, Indicates the first The complex weight matrix corresponding to each frequency band This represents the Hadamard element-wise product; through multi-band spectral enhancement, the model can strengthen macro-environmental patterns such as road grids, building cluster orientations, terrain trends, and repeating textures.

[0049] Step S3033: Perform inverse Fourier transform on the enhanced frequency domain features to reconstruct the frequency domain enhancement result back to the spatial domain, and obtain the environmental field branch output features through residual connection and projection mapping. The reconstruction process is expressed as follows: , in, Indicates the inverse Fourier transform. This indicates a tensor shape reconstruction operation. This represents a multilayer perceptron mapping; in some implementations, the environment field branch further refines the reconstructed spatial features locally by combining window attention or shifted window attention, in order to balance global layout perception with local spatial details. The proposed environment field branch modeling principle is as follows: Figure 3 As shown.

[0050] Step S3034: Define the output of the environmental field branch as an environmental field feature token, which is used to implicitly represent the continuous environmental field semantics in the UAV image, including the distribution of roads, terrain, green space, building clusters, and scene layout information formed by repeated textures; the environmental field features and the landmark map features complement each other at the semantic level, wherein the former provides global scene context and the latter provides discrete landmark anchors directly related to the text, thus jointly constituting a graph-field dual-branch parallel modeling framework.

[0051] Step S40 specifically includes: Step S401: Perform bidirectional cross-attention fusion on the landmark map features output by the landmark map branch and the environmental field features output by the environmental field branch to construct a geometry-guided bidirectional cross-attention fusion module, Cross-Attention Fusion (CAF). The CAF module is used to maintain the fine-grained discrete target recognition capability of the landmark map branch while introducing the global layout context represented by the environmental field branch. Through bidirectional information interaction, it achieves landmark saliency enhancement and environmental region focusing, thereby forming a unified cross-modal fusion representation suitable for natural language-guided UAV target retrieval and localization tasks. The principle of the CAF module is as follows: Figure 4 As shown.

[0052] The specific steps for constructing the CAF module include: Step S4011: Set the output of the landmark map branch to a landmark token sequence. The environment field branch output is an environment token sequence. Considering the high spatial resolution of the environment field branch, to reduce the computational complexity of cross-attention, a representative subset is first uniformly sampled from the environment token sequence. ,in, This represents a subset of environment tokens that participate in the cross-attention calculation.

[0053] Step S4012: To maintain spatial consistency during graph-field fusion, a geometric bias term is introduced in the calculation of cross-attention weights. This allows the model to prioritize context regions that are physically closer to the current target; where, the first The first landmark node and the first The geometric offset between environment tokens is defined as follows: , in, Indicates the first Normalized spatial coordinates of each landmark node Indicates the first The normalized spatial coordinates of an environment token. This represents the learnable distance attenuation coefficient; the geometric bias term is used to penalize long-distance interactions, making the fusion process more consistent with the local spatial consistency constraints in the UAV perspective scene.

[0054] Step S4013: In the CAF module, the context injection process for the landmark query environment is first executed, i.e., Landmark-to-Environment or... In this stage, the landmark token output by the landmark map branch serves as the query, and the environment token output by the environment field branch serves as the key and value. This allows discrete landmarks to absorb relevant global contextual information, thereby enhancing target discrimination capabilities. Its update form is represented as follows: , in, This represents the query matrix obtained by linear mapping of landmark tokens. and These represent the key matrix and value matrix obtained by linear mapping from the environment token, respectively. This represents the channel dimension of the attention head.

[0055] Step S4014: In the CAF module, the target location process for the environment query landmark is further executed, i.e., Environment-to-Landmark or Phase 1; In this phase, the environment token output by the environment field branch serves as the query, and the landmark token output by the landmark map branch serves as the key and value. This allows the environment field representation to reverse-focus and reweight salient target areas based on the distribution of key landmarks. Its update form is represented as follows: , in, This represents the query matrix obtained by linear mapping of the environment token. and These represent the key matrix and value matrix obtained by linear mapping of landmark tokens, respectively; through the reverse aggregation process, the environmental field features are further transformed from a global background description into a contextualized background representation constrained by key targets.

[0056] Step S4015: Update the landmark features obtained in step S4013. The updated environment features obtained in step 4.2.4 By performing stitching, projection mapping, or multilayer perceptron transformation, the image-field fusion feature representation is obtained. The graph-field fusion feature includes both the local structural information of discrete targets and the global layout information of continuous environmental fields, which are used for subsequent graph-text matching, target localization and multi-view consistent reasoning. The bidirectional interaction mechanism is different from the one-way injection fusion strategy and can maintain the precision of point target representation and the stability of field layout representation at the same time.

[0057] Step S402: Based on the graph-field fusion feature representation, jointly train the entire dual-branch parallel modeling framework to construct the total loss function. This ensures that the model simultaneously meets the requirements of global graph-text semantic alignment, target localization accuracy, and graph structure semantic consistency; the total loss function is expressed as... , in, This indicates the learning loss due to the comparison between text and images. This represents the image-text matching loss. This represents the target bounding box localization loss. This represents the graph structure alignment loss. and This represents the weighting coefficient of the corresponding loss term.

[0058] The target bounding box localization loss A summation of bounding box regression loss and generalized intersection-union (OCU) loss is used to constrain the localization accuracy of the landmark map branch for the text-indicated target. Its expression is: , in, Indicates the predicted bounding box. Represents the actual bounding box.

[0059] The graph structure alignment loss This is used to constrain the consistency between visual and text graphs in terms of node semantic distribution, spatial relationships, node discriminativeness, and structural dependencies. Its expression is: , in, , , and These are the weighting coefficients for the corresponding loss terms; the cross-modal distribution alignment loss The expression used to narrow the difference between text node embedding distribution and visual node embedding distribution is: , in, and These represent the probability distributions of text node embedding and visual node embedding, respectively.

[0060] Spatial Relationship Regression Loss Used to constrain the consistency between the predicted spatial relationship vectors between nodes in the visual graph and their true relative coordinates, its expression is: , in, Represents the predicted spatial relationship vector. Represents true relative coordinates, This indicates the number of edges in the visual graph.

[0061] Node-level comparison loss This is used to enhance the discriminative ability of semantically matching node pairs, making positively matching node pairs closer in the embedding space. Its expression is: , in, This represents the set of positively matching node pairs. Indicates the temperature coefficient. This represents the similarity function.

[0062] Structural consistency loss The expression used to constrain the consistency between the text graph dependency matrix and the visual graph adjacency matrix is ​​as follows: , in, Represents the text graph dependency matrix. Represents the adjacency matrix of the visual graph. Represents the Frobenius norm; through the stated This enables the landmark map branches to have stronger cross-modal structural alignment capability before fusion.

[0063] Step S403: End-to-end joint optimization is performed on the visual encoder, text encoder, landmark map branch, environmental field branch, and CAF fusion module using the total loss function to obtain the graph-field bidirectional fusion strategy model GFDFS for natural language-guided UAV target retrieval and localization tasks. Specifically, when only unidirectional cross-attention is retained or the outputs of the two branches are directly concatenated, the model performance is lower than that of the bidirectional CAF fusion method; while simultaneously retaining... and When there is cross-attention in two directions, the image and text retrieval performance reaches its best. The bidirectional cross-attention fusion mechanism can effectively learn the correlation between environmental field features and landmark map features, and improve the final cross-modal retrieval and localization performance.

[0064] The beneficial effects of the natural language-guided UAV target retrieval and localization method based on graph-field bidirectional fusion of the present invention are: (1) This invention uses a dual-path collaborative modeling mechanism of “landmark map + environment field” to explicitly depict the topological relationship between discrete targets and implicitly depict the layout information of continuous scene. Compared with the scheme of using only global features or only local areas, it can more accurately respond to the targets, reference objects and spatial relationships involved in natural language.

[0065] (2) The present invention adopts a bidirectional cross-attention fusion mechanism with geometric bias, which enables local landmarks to actively index the global background context, and enables the global environment field to focus on key landmarks in reverse, thereby effectively suppressing background noise in wide-area aerial images and enhancing the ability to distinguish small targets.

[0066] (3) The present invention adopts a joint optimization method that combines cross-modal comparison, cross-modal matching, localization supervision and graph structure alignment, which is conducive to improving semantic alignment accuracy, spatial reasoning ability and target localization accuracy at the same time. It is applicable to various application scenarios such as natural language guided UAV retrieval, UAV navigation, target matching and geolocation.

[0067] To achieve the above objectives, this invention also proposes a natural language-guided UAV target retrieval and localization system that integrates graph and field representations, such as... Figure 6As shown, the system includes a processor 1001, a CPU, a network interface 1004, a user interface 1003, a memory 1005, a communication bus 1002, and a graph-field bidirectional fusion natural language-guided UAV target retrieval and localization program stored on the processor. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen or an input unit such as a keyboard. Optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be a high-speed RAM or a stable, non-volatile memory, such as a disk drive. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0068] Those skilled in the art will understand that Figure 6 The system structure shown does not constitute a limitation on the system and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0069] like Figure 6 As shown, the memory 1005, which serves as a computer storage medium, may include an operating device, a network communication module, a user interface module, and a natural language-guided UAV target retrieval and localization program that integrates image and field fusion.

[0070] exist Figure 6 In the system shown, the network interface 1004 is mainly used to connect to the network server and communicate with the network server; the user interface 1003 is mainly used to interact with the user terminal and receive user input commands; and the processor 1001 can be used to call the graph-field bidirectional fusion natural language guided UAV target retrieval and positioning program stored in the memory 1005.

[0071] To achieve the above objectives, the present invention also proposes a computer-readable storage medium storing a graph-field bidirectional fusion natural language-guided UAV target retrieval and localization program. When the graph-field bidirectional fusion natural language-guided UAV target retrieval and localization program is run by a processor, the steps of the method described above are executed, which will not be repeated here.

[0072] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.

Claims

1. A natural language-guided target retrieval and localization method for unmanned aerial vehicles (UAVs) using a two-way graph-field fusion model, characterized in that, The method includes the following steps: Step S10: A theoretical analysis is conducted on the cross-modal and multi-perspective complex challenges in natural language-guided UAV target retrieval and localization tasks. Step 20: Construct a graph-field decoupling and collaborative modeling mechanism to prove that combining discrete landmark relationship modeling with global environment layout modeling can optimize the cross-modal, multi-view composite problem; Step S30: Establish a graph-field dual-branch parallel modeling framework for natural language guided UAV missions, specifically including: constructing a landmark map branch using a map salient target display modeling method, and constructing an environmental field branch using an environmental macro-layout implicit modeling method; Step 40: Train the graph-field dual-branch parallel modeling framework, and use a bidirectional cross-attention fusion mechanism to jointly optimize the landmark graph branch and the environmental field branch to obtain a graph-field bidirectional fusion strategy model for natural language-guided UAV target retrieval and localization. Step S50: Train the trained graph-field bidirectional fusion strategy model on the test set and output the target retrieval results and / or target localization results guided by natural language.

2. The graph-field bidirectional fusion natural language-guided UAV target retrieval and localization method according to claim 1, characterized in that, Step S10 includes: Step S101: Define the natural language-guided UAV target retrieval and localization task as a cross-modal semantic alignment task based on UAV viewpoint images, with the input being the UAV viewpoint image. With natural language instructions Among them, the images from the perspective of the drone satisfy Natural language instructions It includes target entities, attribute descriptions, and spatial relationship semantics; the goal of the localization task is to locate objects in a visual dataset. With text datasets The above establishes a bidirectional cross-modal correspondence between images and text to achieve text-guided target retrieval, image-guided text retrieval, and target localization. Step S102 defines the natural language-guided UAV target retrieval and localization task as a unified retrieval and localization problem. The retrieval subtask retrieves the most relevant image from a candidate UAV image set based on natural language instructions, or queries the most relevant text based on an image. The localization subtask determines the location of the target region in the image based on natural language instructions. The retrieval subtask learns a graph-text similarity function in a joint embedding space. The result of the search is expressed as follows: , in, Represents a visual dataset. Represents a text dataset. Represents a visual encoder. Indicates a text encoder; Step S103, define the localization subtask as: for natural language instructions The target region is pointed to, and its prediction in the image is made. bounding box in ,in, Indicates the center coordinates of the target area. and These represent the width and height of the bounding box, respectively; the bounding box is predicted by solving the equation. With the true bounding box The deviation between them is used to achieve natural language-guided target localization, and the localization loss is expressed as: ; Step S104 further defines the natural language-guided UAV task as a cross-modal, multi-view coupling problem. Cross-modal coupling means that the local target semantics in the natural language description needs to establish an accurate correspondence with the local area in the wide-area aerial image. Multi-view coupling means that the same target has significant representational shifts under different imaging angles, scales, and background conditions. Therefore, the core of the localization subtask is to establish an accurate correspondence between discrete landmarks and combined semantics in a large-scale environmental background. Step S105: Based on the definition of the localization subtask, it is clarified that the method needs to simultaneously meet two types of modeling requirements: local salient target recognition and global environment layout perception. Local salient target recognition is used to enhance the sensitivity to discrete targets indicated by the text and their spatial relationships, while global environment layout perception is used to provide the macroscopic context of the scene where the target is located, thereby providing a problem definition basis for subsequent graph-field dual-branch parallel modeling.

3. The graph-field bidirectional fusion natural language-guided UAV target retrieval and localization method according to claim 2, characterized in that, Step S20 includes: Step S201 addresses the cross-modal granularity mismatch problem in natural language-guided UAV target retrieval and localization tasks, where "text only covers local target semantics, while images contain a large range of complex environmental backgrounds." A graph-field decoupling and collaborative modeling mechanism is constructed, decomposing visual representation into two parallel pathways: discrete landmark relationship modeling and global environment layout modeling. Discrete landmark relationship modeling captures the local targets referred to in the text and their spatial topological relationships, while global environment layout modeling captures the macro-context formed by the distribution of roads, terrain, building clusters, and repetitive textures. Through bidirectional interactive fusion, the two types of representations form complementary constraints within a unified embedding space, thereby suppressing background interference and enhancing the semantic saliency of small targets. Step S202 demonstrates how discrete landmark relationship modeling improves model performance, specifically including: Step S021: Establish a local alignment propagation theory for the discrete landmark relationship modeling branch, defining the visual graph and text graph in the first step. The node embeddings of the layers are respectively and The final fused visual representation is defined as follows: , in, This represents the output features of the visual backbone network. Indicate the number of propagation layers in the graph neural network; and assume the message passing function. Fusion function with cross attention It satisfies the Lipschitz continuity condition, that is: , , It is also assumed that the graph topological differences are constrained by an upper bound on the semantic differences of the nodes, that is: , in, This represents the adjacency structure of the visual graph. This represents the adjacency structure of a text graph. Indicates the structural coupling coefficient; Step S2022, under the assumptions of step S2021, proves that the local node-level alignment error can propagate along the inter-layer paths of the graph neural network and ultimately constrain the global cross-modal representation differences; wherein, the initial node-level alignment error is defined as: , The final cross-modal representation differences satisfy: , in, Represents visual backbone feature noise. This indicates irreducible modal differences; thus, it shows that when node-level semantic alignment error... During descent, the final fusion representation and The cross-modal deviation is reduced synchronously, which theoretically proves that discrete landmark relationship modeling can propagate local semantic alignment into global representation consistency. Step S2023, further prove the inter-layer error propagation in step S2022, and define the first... The differences in layer graph node embedding are: , Then, from the decomposition of the characteristic error term and the structural error term, we can obtain: , Combining the structural consistency assumption, we can further obtain: , We can obtain the following through recursion: , Substituting the Lipschitz upper bound of the cross-attention fusion function, we obtain the final cross-modal representation difference upper bound, thus completing the theoretical proof of the discrete landmark relationship modeling branch; Step S203 demonstrates how global environment layout modeling improves model performance, specifically including: Step S2031: Establish frequency domain alignment propagation theory for the global environment layout modeling branch, and define the visual backbone output as... Its frequency domain representation is defined as: ,in, Represent the Discrete Fourier Transform; let the adaptive frequency mixing operator be... The modulated visual spectrum is then: And the environmental field representation is reconstructed through inverse transformation: Let the ideal environment field, implicitly defined by the text global vector, be represented as... The corresponding ideal spectrum is: The theoretical goal of the environmental field branch is to make and Alignment; Step S2032: Establish the frequency domain and spatial domain error equivalence relationship for step S2031, assuming that the discrete Fourier transform satisfies Parseval's theorem, that is, for any signal... and ,have: , And it is assumed that the frequency mixing operator satisfies Lipschitz continuity: , Meanwhile, assume that the spectrum alignment residual satisfies: , in, Indicates frequency domain alignment error; Step S2033, under the assumptions of step S2032, prove that frequency domain alignment can propagate into spatial domain environmental field uniformity; let The spatial inconsistency of the global environment field satisfies: , in, This represents the approximation error of the ideal environmental field; further considering the visual backbone input noise, we have: , in, This represents the ideal visual input; thus proving that the frequency domain alignment error... The optimization will linearly improve the consistency of the spatial domain environment field, theoretically demonstrating that global environment layout modeling can effectively enhance the robustness of cross-modal alignment; Step S204: Based on the theoretical results of steps S202 and S203, the inference of the graph-field decoupling and collaborative modeling mechanism is obtained: discrete landmark relationship modeling is responsible for propagating node-level semantic alignment error constraints to the global fusion representation, improving target-level and relationship-level identification capabilities; global environment layout modeling is responsible for propagating spectral alignment errors into spatial field consistency, improving scene-level and context-level robustness.

4. The graph-field bidirectional fusion natural language-guided UAV target retrieval and localization method according to claim 3, characterized in that, Step S30 includes: Step S301: Construct a graph-field dual-branch parallel modeling framework. Perform local salient target modeling and global environment layout modeling on the input UAV view image and natural language commands, respectively. The local salient target modeling forms a landmark map branch; the global environment layout modeling forms an environment field branch. The landmark map branch is used to explicitly represent target entities and their spatial topological relationships, and the environment field branch is used to implicitly represent scene texture distribution, repeating structures and macroscopic layout information. The two branches are processed in parallel and output landmark map features and environment field features respectively. The specific steps for constructing the landmark map branch include: Step S3021: Extract the target from the UAV image using a text-guided target detector in the visual domain. Identify candidate salient landmark areas and construct a visual graph using the bounding boxes of each candidate area as nodes. in, Represents the set of nodes in a visual graph. Represents the set of edges in a visual graph; to characterize the relative spatial relationships between landmarks, nodes are... With nodes The relative center position and scale information between them are encoded into the edge attributes, and its spatial edge features are defined as follows: , in, and Representing nodes respectively With nodes The center coordinates of the corresponding bounding box These represent the width and height of the corresponding bounding box, respectively; through the above spatial edge features, explicit modeling of direction, distance, and scale changes in the visual landmark map is achieved. Step S3022: In the landmark graph branch, graph inference based on edge attribute enhancement is performed on the visual graph nodes. The TransformerConv operator is used to aggregate the features of neighboring nodes and edge space features, so that each visual node obtains the topological context related to its neighboring landmarks. The update rule is as follows: , in, Indicates the first Layer nodes Visual features Represents a node The neighborhood set, Represents a node For nodes Attention weights , and The learnable parameter matrix is ​​represented; the update rule enables visual graph nodes to retain explicit spatial geometric constraints while aggregating neighborhood information. Step S3023: Construct a text graph corresponding to the visual graph in the text field, denoted as... ,in, Represents a set of text graph nodes. This represents the text graph edge set. A large language model is used to extract landmark entities, attribute modifiers, and relational descriptions from natural language instructions to initialize text node features. Directional encoding is also used to initialize text graph edge features to characterize the syntactic dependencies and spatial relationships between different semantic entities. Furthermore, a graph attention network is employed to perform semantic propagation on the text graph, with attention coefficients and node update rules as follows: , , , in, Indicates the first Layer nodes Textual features, This represents vector concatenation. and Indicates learnable parameters, It represents a non-linear activation function; through the text graph construction and update process, phrase-level semantics are aggregated into a contextualized landmark semantic representation containing target, attribute, and relation information; Step S3024: Align and represent the visual node features after visual graph inference with the text node features after text graph propagation to obtain the landmark map feature token output by the landmark map branch. This token is used to explicitly characterize the discrete targets corresponding to the natural language instructions, the spatial relationships between targets, and the local geometric topology. The landmark map features are used to enhance the model's ability to recognize spatial relationship descriptions such as center, left, right, top, bottom, and adjacency, and to improve the distinguishability between small targets and non-dominant targets. The specific steps for constructing the environmental field branch include: Step S3031: Process the global feature map output by the visual backbone network. Perform serialization processing on the input feature map Flattened to a length of sequence representation It is then mapped to the frequency domain using a one-dimensional real discrete Fourier transform, and its frequency domain representation is as follows: , in, Represents the imaginary unit. It represents the frequency index; by transforming a two-dimensional spatial layout into a one-dimensional frequency domain sequence, it enables a global perception of wide-area environmental characteristics. Step S3032, in the environmental field branch, to adaptively enhance the frequency band related to the environmental layout, generate... A learnable one-dimensional frequency band mask is used, and weighted modulation is performed on each frequency band; where, the first... A frequency band mask is defined as: , in, Indicates the first The center frequency of each frequency band Indicates the first The bandwidth parameters of each frequency band are determined; further, the frequency domain features are modulated element-wise with the frequency band mask and the learnable complex weight matrix to obtain the enhanced frequency domain features: , in, Indicates the first The complex weight matrix corresponding to each frequency band This represents the Hadamard element-wise product; through multi-band spectral enhancement, the model can strengthen macro-environmental patterns such as road grids, building cluster orientations, terrain trends, and repetitive textures. Step S3033: Perform inverse Fourier transform on the enhanced frequency domain features to reconstruct the frequency domain enhancement result back to the spatial domain, and obtain the environmental field branch output features through residual connection and projection mapping. The reconstruction process is expressed as follows: , in, Indicates the inverse Fourier transform. This indicates a tensor shape reconstruction operation. This represents a multilayer perceptron mapping; in some implementations, the environmental field branch further refines the reconstructed spatial features by combining window attention or shifted window attention, in order to take into account both global layout perception and local spatial details. Step S3034: Define the output of the environmental field branch as an environmental field feature token, which is used to implicitly represent the continuous environmental field semantics in the UAV image, including the distribution of roads, terrain, green space, building clusters, and scene layout information formed by repeated textures; the environmental field features and the landmark map features complement each other at the semantic level, wherein the former provides global scene context and the latter provides discrete landmark anchors directly related to the text, thus jointly constituting a graph-field dual-branch parallel modeling framework.

5. The graph-field bidirectional fusion natural language-guided UAV target retrieval and localization method according to claim 4, characterized in that, Step S40 includes: Step S401: Perform bidirectional cross-attention fusion on the landmark map features output by the landmark map branch and the environmental field features output by the environmental field branch to construct a geometry-guided bidirectional cross-attention fusion module; wherein, the bidirectional cross-attention fusion module is used to introduce the global layout context represented by the environmental field branch while maintaining the fine-grained discrete target recognition capability of the landmark map branch, and realize landmark saliency enhancement and environmental region focusing through bidirectional information interaction, thereby forming a unified cross-modal fusion representation suitable for natural language-guided UAV target retrieval and localization tasks; The specific steps for constructing the bidirectional cross-attention fusion module include: Step S4011: Set the output of the landmark map branch to a landmark token sequence. The environment field branch output is an environment token sequence. Considering the high spatial resolution of the environment field branch, to reduce the computational complexity of cross-attention, a representative subset is first uniformly sampled from the environment token sequence. ,in, This represents a subset of environment tokens that participate in the cross-attention calculation; Step S4012: To maintain spatial consistency during graph-field fusion, a geometric bias term is introduced in the calculation of cross-attention weights. This allows the model to prioritize context regions that are physically closer to the current target; where, the first The first landmark node and the first The geometric offset between environment tokens is defined as follows: , in, Indicates the first Normalized spatial coordinates of each landmark node Indicates the first The normalized spatial coordinates of an environment token. This represents the learnable distance attenuation coefficient; the geometric bias term is used to penalize long-distance interactions, making the fusion process more consistent with the local spatial consistency constraints in the UAV perspective scene. Step S4013: In the bidirectional cross-attention fusion module, the context injection process for the landmark query environment is first executed, i.e., Landmark-to-Environment or... In this stage, the landmark token output by the landmark map branch serves as the query, and the environment token output by the environment field branch serves as the key and value. This allows discrete landmarks to absorb relevant global contextual information, thereby enhancing target discrimination capabilities. Its update form is expressed as: , in, This represents the query matrix obtained by linear mapping of landmark tokens. and These represent the key matrix and value matrix obtained by linear mapping from the environment token, respectively. The channel dimension representing the attention head; Step S4014: In the bidirectional cross-attention fusion module, the target location process of the environmental query landmark is further executed. In this process, the environmental token output by the environmental field branch is used as the Query, and the landmark token output by the landmark map branch is used as the Key and Value. This enables the environmental field representation to reverse focus and reweight the significant target area based on the distribution of key landmarks. Its update form is expressed as follows: , in, This represents the query matrix obtained by linear mapping of the environment token. and These represent the key matrix and value matrix obtained by linear mapping of landmark tokens, respectively; through the reverse aggregation process, the environmental field features are further transformed from a global background description into a contextualized background representation constrained by key targets; Step S4015: Update the landmark features obtained in step S4013. The updated environment features obtained in step S4014 By performing stitching, projection mapping, or multilayer perceptron transformation, the image-field fusion feature representation is obtained. The graph-field fusion feature contains both the local structural information of discrete targets and the global layout information of continuous environmental fields, which are used for subsequent graph-text matching, target localization, and multi-view consistent reasoning. The bidirectional interaction mechanism differs from the unidirectional injection fusion strategy and can maintain both the fineness of point target representation and the stability of field layout representation. Step S402: Based on the graph-field fusion feature representation, jointly train the entire dual-branch parallel modeling framework to construct the total loss function. This enables the model to simultaneously meet the requirements of global graph-text semantic alignment, target localization accuracy, and graph structure semantic consistency; the total loss function is expressed as: , in, This indicates the learning loss due to the comparison between text and images. This represents the image-text matching loss. This represents the target bounding box localization loss. This represents the graph structure alignment loss. and This represents the weighting coefficient of the corresponding loss term; The target bounding box localization loss The sum of the bounding box regression loss and the generalized intersection-union (OCU) loss is used to constrain the localization accuracy of the landmark map branch for the text-indicated target. Its expression is: , in, Indicates the predicted bounding box. Represents the true bounding box; The graph structure alignment loss The expression used to constrain the consistency between visual and text graphs in terms of node semantic distribution, spatial relationships, node discriminativeness, and structural dependencies is: , in, , , and These are the weighting coefficients for the corresponding loss terms; the cross-modal distribution alignment loss The expression used to narrow the difference between text node embedding distribution and visual node embedding distribution is: , in, and These represent the probability distributions of text node embedding and visual node embedding, respectively. Spatial Relationship Regression Loss The expression used to constrain the consistency between the predicted spatial relationship vectors between nodes in the visual graph and their true relative coordinates is: , in, Represents the predicted spatial relationship vector. Represents true relative coordinates, Indicates the number of edges in the visual graph; Node-level comparison loss To enhance the discriminative power of semantically matching node pairs, making positively matching node pairs closer in the embedding space, its expression is: , in, This represents the set of positively matching node pairs. Indicates the temperature coefficient. Represents the similarity function; Structural consistency loss The expression used to constrain the consistency between the text graph dependency matrix and the visual graph adjacency matrix is: , in, Represents the text graph dependency matrix. Represents the adjacency matrix of the visual graph. Represents the Frobenius norm; through the stated This enables the landmark map branches to have stronger cross-modal structural alignment capability before fusion; Step S403: The visual encoder, text encoder, landmark map branch, environmental field branch, and bidirectional cross-attention fusion module are jointly optimized end-to-end through the total loss function to obtain a graph-field bidirectional fusion strategy model for natural language-guided UAV target retrieval and localization tasks.

6. The graph-field bidirectional fusion natural language-guided UAV target retrieval and localization method according to claim 5, characterized in that, Step S50 includes: inputting the test set into the pre-trained model, and the model outputting the predicted answer.

7. A natural language-guided UAV target retrieval and localization system that integrates graph and field data, characterized in that... The system includes a memory, a processor, and a graph-field bidirectional fusion natural language guided UAV target retrieval and localization program stored on the processor, wherein the graph-field bidirectional fusion natural language guided UAV target retrieval and localization program is executed by the processor to perform the steps of the method as described in any one of claims 1 to 6.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a graph-field bidirectional fusion natural language-guided UAV target retrieval and localization program, which, when run by a processor, performs the steps of the method as described in any one of claims 1 to 6.