Fine alignment learning navigation model training method, navigation method and device

By fusing visual, semantic, and spatial features from UAV aerial images and utilizing multiple auxiliary prediction tasks for refined alignment learning, an aerial visual dialogue navigation model is generated. This solves the problem of inaccurate UAV navigation in complex scenarios and achieves efficient and accurate navigation path planning.

CN120014488BActive Publication Date: 2025-12-09INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510043915.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-12-09
Estimated Expiration
2045-01-10

AI Technical Summary

Technical Problem

Existing UAV navigation methods rely on coarse-grained global feature matching, which makes it difficult to accurately understand the semantic descriptions and positional relationships of multiple landmarks in complex scenarios. This results in inaccurate navigation predictions and makes it difficult to capture the relationships between entity attributes in the language and landmarks, leading to inaccurate navigation path planning.

Method used

By extracting visual, semantic, and spatial features from drone aerial images, weighted fusion is performed to generate semantic mesh features. Multiple auxiliary prediction tasks are used for refined alignment learning, including landmark rotation bounding box prediction, landmark semantic prediction, and entity-landmark contrast learning. Combined with a refined aerial visual dialogue navigation dataset, iterative training is performed to generate an aerial visual dialogue navigation model.

Benefits of technology

It improves the navigation accuracy and alignment capability of UAVs in complex scenarios, enhances the efficiency and accuracy of navigation planning, and enables real-time interaction and decision-making under multimodal information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014488B_ABST
    Figure CN120014488B_ABST
Patent Text Reader

Abstract

The application provides a navigation model training method, a navigation method and a device for fine alignment learning, and the navigation model training method comprises the following steps: weighting and fusing visual features, semantic features and spatial features extracted from unmanned aerial vehicle aerial images to obtain semantic grid features; performing display learning on fine alignment relationships between entities and landmark objects based on the semantic grid features according to multiple auxiliary prediction tasks to obtain visual representations; taking a fine aerial visual dialogue navigation data set as a training sample, taking the visual representations as input features, and taking a comprehensive loss as a loss function to iteratively train a navigation model to obtain an aerial visual dialogue navigation model; wherein the comprehensive loss is determined based on a navigation loss function and loss functions corresponding to the multiple auxiliary prediction tasks. The method comprehensively fuses multi-modal features, and improves the navigation accuracy, alignment capability and task execution efficiency of the unmanned aerial vehicle in a complex scene.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer vision, in particular to a navigation model training method based on fine alignment learning, a navigation method and device. BACKGROUND

[0002] In recent years, multi-modal learning has gradually become an important research direction in the field of artificial intelligence, and its core goal is to improve the understanding and reasoning ability of the model in complex tasks by fusing visual and language information; in this context, the aerial vision-dialog navigation task (Aerial Vision-Dialog Navigation, AVDN) as a complex task combining visual perception, language understanding and navigation planning has received extensive attention.

[0003] In related technologies, most of the existing unmanned aerial vehicle navigation methods rely on coarse-grained global feature matching, and lack fine-grained capture of visual features and language semantics, especially in complex scenes, the understanding ability of the semantic description and position relationship of multiple landmarks is weak, resulting in inaccurate unmanned aerial vehicle navigation prediction results; on the other hand, when processing complex instructions, the unmanned aerial vehicle is difficult to capture the combined information of entity attributes, relationships between landmarks and priorities in language at the same time, and due to the lack of fine-grained modeling of landmark semantics, especially in dynamic navigation tasks, it is difficult to realize real-time interaction and decision-making of multi-modal information, resulting in inaccurate navigation path planning. SUMMARY

[0004] The present application provides a navigation model training method based on fine alignment learning, a navigation method and device, to solve the defects that the existing unmanned aerial vehicle navigation method relies on coarse-grained global feature matching, the understanding ability of the semantic description and position relationship of multiple landmarks is weak in complex scenes, resulting in inaccurate unmanned aerial vehicle navigation prediction results, and the existing technology is difficult to capture the combined information of entity attributes, relationships between landmarks and priorities in language at the same time, resulting in inaccurate navigation path planning, improve the unmanned aerial vehicle navigation planning efficiency and accuracy.

[0005] The present application provides a navigation model training method based on fine alignment learning, comprising:

[0006] Extracting visual features, semantic features and spatial features from the unmanned aerial vehicle aerial image, and weighting and fusing the visual features, the semantic features and the spatial features to obtain semantic grid features; wherein the semantic features are used to represent the class information and geometric information of the landmark objects in the unmanned aerial vehicle aerial image; the spatial features are used to represent the relative direction and relative distance between the unmanned aerial vehicle and the landmark objects;

[0007] display learning of a fine-grained alignment relationship between entities and landmark objects according to the semantic grid features based on a plurality of auxiliary prediction tasks, to obtain visual representations; wherein the plurality of auxiliary prediction tasks include at least two of a landmark rotation bounding box prediction task, a landmark semantic prediction task, and an entity-landmark contrastive learning task;

[0008] Taking a fine-grained aerial visual dialog navigation dataset as a training sample, taking the visual representations as input features, and taking a comprehensive loss as a loss function, a navigation model is iteratively trained to obtain an aerial visual dialog navigation model; wherein the comprehensive loss is determined based on a navigation loss function and a loss function corresponding to the plurality of auxiliary prediction tasks; the fine-grained aerial visual dialog navigation dataset is obtained by fine-grained expansion of entity-landmark data in an unmanned aerial vehicle dialog navigation (AVDN) dataset.

[0009] According to the navigation model training method of fine-grained alignment learning provided by the application, the fine-grained aerial visual dialog navigation dataset is obtained by the following steps:

[0010] Language entities are extracted from navigation dialog data in an AVDN dataset; the language entities include a plurality of landmark objects, landmark orientation descriptions, and semantic relationships;

[0011] Landmark masks are obtained from the AVDN dataset, and rotation bounding box information of the landmark masks is extracted, the rotation bounding box information including a landmark center point, a width, a height, and a rotation angle;

[0012] The language entities and the landmark masks are preliminarily aligned, and entity-landmark pairs that are incorrectly matched are corrected through manual verification to obtain the fine-grained aerial visual dialog navigation data.

[0013] According to the navigation model training method of fine-grained alignment learning provided by the application, the plurality of auxiliary prediction tasks include the landmark rotation bounding box prediction task, the landmark semantic prediction task, and the entity-landmark contrastive learning task;

[0014] The display learning of the fine-grained alignment relationship between entities and landmark objects according to the semantic grid features based on the plurality of auxiliary prediction tasks includes:

[0015] Landmark rotation bounding box information is obtained through the landmark rotation bounding box prediction task based on a multi-head cross-attention mechanism and the semantic grid features;

[0016] Landmark language description information is obtained through the landmark semantic prediction task based on a self-recurrent mechanism and the semantic grid features;

[0017] The cross-modal features of the entity and the landmark are subjected to contrastive learning according to the landmark rotation bounding box information and the landmark language description information through the entity-landmark contrastive learning task, so as to obtain the visual representation.

[0018] According to the navigation model training method for fine alignment learning provided by the application, the comprehensive loss is determined by the following formula:

[0019] ;

[0020] Among them, the comprehensive loss, the navigation loss function, the loss function corresponding to the plurality of auxiliary prediction tasks;

[0021] It is expressed by the following formula:

[0022] ;

[0023] Among them, the loss function corresponding to the landmark rotation bounding box prediction task, the loss function corresponding to the landmark semantic prediction task, the loss function corresponding to the entity-landmark contrastive learning task; and is a weight parameter for balancing the loss weight of each task.

[0024] According to the navigation model training method for fine alignment learning provided by the application, after obtaining the aerial visual dialogue navigation model, the method further comprises:

[0025] The navigation strategy model based on time sequence transformation fuses and decides the navigation result output by the aerial visual dialogue navigation model to obtain a target navigation strategy; wherein the navigation strategy model based on time sequence transformation is used to obtain the semantic grid feature, the existing dialogue data, the visual feature and the trajectory data and the predicted flight trajectory of the unmanned aerial vehicle.

[0026] The application also provides a navigation method, comprising:

[0027] Obtaining an unmanned aerial vehicle aerial image to be processed;

[0028] Processing the unmanned aerial vehicle aerial image to be processed based on an aerial visual dialogue navigation model to obtain a navigation result; the aerial visual dialogue navigation model is obtained by training the navigation model training method for fine alignment learning;

[0029] The navigation strategy model based on time sequence transformation fuses and decides the navigation result to obtain a navigation strategy, wherein the navigation strategy model based on time sequence transformation is used to obtain a navigation trajectory of the unmanned aerial vehicle according to the semantic grid feature, the existing dialogue data, the visual feature, the trajectory data and the predicted navigation trajectory.

[0030] The application further provides a navigation model training device for fine alignment learning, comprising:

[0031] The feature fusion module is used to extract visual features, semantic features and spatial features from the unmanned aerial vehicle aerial image, and to obtain semantic grid features by weighted fusion of the visual features, the semantic features and the spatial features, wherein the semantic features are used to represent class information and geometric information of landmark objects in the unmanned aerial vehicle aerial image, and the spatial features are used to represent relative directions and relative distances between the unmanned aerial vehicle and the landmark objects.

[0032] The alignment learning module is used to display learning of fine alignment relationships between entities and landmark objects based on the semantic grid features according to a plurality of auxiliary prediction tasks to obtain visual representations, wherein the plurality of auxiliary prediction tasks include at least two of landmark rotation bounding box prediction tasks, landmark semantic prediction tasks and entity-landmark contrast learning tasks.

[0033] The training module is used to take a fine aerial visual dialogue navigation data set as a training sample, take the visual representations as input features, and take a comprehensive loss as a loss function to iteratively train a navigation model to obtain an aerial visual dialogue navigation model, wherein the comprehensive loss is determined based on a navigation loss function and loss functions corresponding to the plurality of auxiliary prediction tasks, and the fine aerial visual dialogue navigation data set is obtained by fine-grained expansion of entity-landmark data in an unmanned aerial vehicle dialogue navigation (AVDN) data set.

[0034] The application further provides a navigation device, comprising:

[0035] The image acquisition module is used to acquire unmanned aerial vehicle aerial images to be processed.

[0036] The navigation prediction module is used to process the unmanned aerial vehicle aerial images to be processed based on an aerial visual dialogue navigation model to obtain navigation results, wherein the aerial visual dialogue navigation model is obtained by the fine alignment learning navigation model training method.

[0037] The navigation decision module is used to fuse and decide the navigation results based on a navigation strategy model based on time sequence transformation to obtain a navigation strategy, wherein the navigation strategy model based on time sequence transformation is used to obtain a navigation trajectory of the unmanned aerial vehicle according to the semantic grid feature, the existing dialogue data, the visual feature, the trajectory data and the predicted navigation trajectory.

[0038] The application further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the navigation model training method or the navigation method of the fine alignment learning according to any one of the above when executing the computer program.

[0039] The application further provides a non-transitory computer-readable storage medium, which stores a computer program, wherein the computer program is executable on a processor to implement the navigation model training method or the navigation method of the fine alignment learning according to any one of the above.

[0040] The application further provides a computer program product, comprising a computer program, wherein the computer program is executable on a processor to implement the navigation model training method or the navigation method of the fine alignment learning according to any one of the above.

[0041] The navigation model training method, the navigation method and the device of the fine alignment learning provided by the application can obtain semantic grid features by weighting and fusing the visual features, the semantic features and the spatial features extracted from the aerial images of the unmanned aerial vehicle, can display and learn the fine alignment relationship between the entities and the landmark objects according to the semantic grid features by using multiple auxiliary prediction tasks, can obtain visual representations, and finally can obtain an aerial visual dialogue navigation model by taking the fine aerial visual dialogue navigation dataset as a training sample, taking the visual representations as input features, and iteratively training the navigation model by taking a comprehensive loss as a loss function, so that the navigation precision, the alignment capability and the task execution efficiency of the unmanned aerial vehicle in a complex scene are improved by comprehensively fusing the multi-modal features (including visual, semantic and spatial information). BRIEF DESCRIPTION OF DRAWINGS

[0042] In order to more clearly illustrate the technical solutions of the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are some embodiments of the application, and other drawings can be obtained by those skilled in the art without any creative effort on the basis of these drawings.

[0043] Figure 1 is one of the flowcharts of the navigation model training method of the fine alignment learning provided by the application.

[0044] Figure 2 is another flowchart of the navigation model training method of the fine alignment learning provided by the application.

[0045] Figure 3 is the flowchart of the navigation method provided by the application.

[0046] Figure 4It is a structural schematic diagram of the navigation model training device for fine alignment learning provided by the application.

[0047] Figure 5 It is a structural schematic diagram of the navigation device provided by the application.

[0048] Figure 6 It is a structural schematic diagram of the electronic device provided by the application. DETAILED DESCRIPTION

[0049] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below with reference to the drawings in the present application. Obviously, the described embodiments are some embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the protection scope of the present application.

[0050] The navigation model training method, navigation method and device for fine alignment learning of the present application will be described below. Figures 1-5 The navigation model training method, navigation method and device for fine alignment learning of the present application will be described below.

[0051] Figure 1 It is one of the flowcharts of the navigation model training method for fine alignment learning provided by the application, as shown in the figure, the navigation model training method comprises the following steps: Figure 1

[0052] Step 110, extracting visual features, semantic features and spatial features from the aerial image of the unmanned aerial vehicle, and performing weighted fusion on the visual features, semantic features and spatial features to obtain semantic grid features; wherein the semantic features are used to represent the category information and geometric information of the landmark object in the aerial image of the unmanned aerial vehicle; the spatial features are used to represent the relative direction and relative distance between the unmanned aerial vehicle and the landmark object.

[0053] In this step, the aerial image of the unmanned aerial vehicle can be obtained from the Aerial Vision-and-Dialog Navigation (AVDN) data, and the Aerial Vision-and-Dialog Navigation (AVDN) data contains corresponding unmanned aerial vehicle-human natural language dialogue.

[0054] In this step, in order to improve the visual perception ability of the navigation model, visual, semantic and spatial features can be extracted from the aerial image of the unmanned aerial vehicle, and semantic grid features can be constructed; the specific process is as follows:

[0055] (1) Extracting visual features from the aerial image of the unmanned aerial vehicle as the preliminary grid features of the scene.

[0056] ​​In this embodiment, the UAV aerial image can be encoded by a visual encoder to obtain corresponding visual features.

[0057] In this embodiment, the visual encoder includes but is not limited to a convolutional neural network (CNN) or a neural network model based on a self-attention mechanism.

[0058] In this embodiment, before feature extraction of the UAV aerial image, the UAV aerial image can be preprocessed, including denoising, enhancement, correction, etc., to improve image quality and reduce the difficulty of subsequent processing.

[0059] (2) The UAV aerial image is detected by a target detection model to generate class labels, boundary information and geometric features of landmark objects, and further generate semantic features of the objects to capture the class and geometric information of the landmark.

[0060] In this embodiment, the target detection model includes but is not limited to the YOLO (You Only Look Once) series and the R-CNN (Regions with Convolutional Neural Networks, region-based convolutional neural network) series, etc.

[0061] In this embodiment, a UAV aerial image dataset is collected and labeled, which can include various landmark objects such as buildings, bridges, sculptures, etc., and label their class labels, boundary boxes and geometric features (such as height, width, area, etc.).

[0062] In this embodiment, the trained target detection model is used to detect the target of the preprocessed image, and the model will identify the landmark object and generate its class label and boundary box. According to the result of target detection, the geometric features of the landmark object are extracted, such as height, width, area, perimeter, etc. These features can be used for subsequent analysis and processing.

[0063] In this embodiment, data labeling can use polygon labeling, boundary box labeling and other methods to ensure the accuracy and integrity of the labeling.

[0064] In this embodiment, the class labels, boundary information and geometric features of the landmark objects can be fused to form a feature vector containing the class information and geometric information of the landmark objects; then the deep learning technology (convolutional neural network or recurrent neural network) is used to further process the feature vector to generate the semantic features of the landmark objects, which are used for subsequent tasks such as landmark recognition, classification, retrieval, etc.

[0065] (3) Combine the location information of the UAV to generate polar coordinate spatial features between the landmark and the UAV. This enhances the model's spatial perception capabilities by describing the relative directions and distances of landmarks.

[0066] In this embodiment, the drone's location information, such as longitude, latitude, and altitude (or height), can be obtained through the drone's positioning device (e.g., GPS system).

[0067] In this embodiment, polar coordinate spatial features are generated based on the location information of the UAV. The formula for representing the relative relationship between a landmark and a drone is:

[0068] ;

[0069] in, The relative angle between the landmark and the drone. The normalized landmark distance; the polar coordinate spatial encoding, by standardizing the spatial distance and direction between the UAV and the target landmark, can explicitly capture the geometric relationship of the navigation path and improve the model's ability to perceive the spatial layout.

[0070] (4) The visual features, semantic encoding and spatial encoding are weighted and fused to generate the final semantic grid features. ; It integrates visual, semantic, and spatial information, and can effectively capture the spatial layout and semantic relationships of landmarks in a scene.

[0071] In this embodiment, semantic grid features The generation is achieved by fusing visual features. Semantic encoding and polar coordinate space encoding The specific integration method is as follows:

[0072] ;

[0073] in, Presentation layer normalization operation, , , These are the learning parameters. The semantic grid features combine visual, semantic, and spatial coding features, enabling UAVs to effectively capture the semantic information and spatial distribution of landmarks in complex scenes.

[0074] In step 120, based on the plurality of auxiliary prediction tasks, display learning of the fine-grained alignment relationship between the entity and the landmark object is performed according to the semantic grid features, and visual representation is obtained. The plurality of auxiliary prediction tasks includes at least two of a landmark rotation bounding box prediction task, a landmark semantic prediction task, and an entity-landmark contrastive learning task.

[0075] In this step, the landmark rotation bounding box prediction task can capture the geometric structure and rotation information of the landmark, the landmark semantic prediction task can obtain the text description information related to the landmark, and the entity-landmark contrastive learning task can improve the alignment accuracy of the navigation model in the multi-modal feature space.

[0076] In this embodiment, in order to enhance the ability of the model in the multi-modal feature alignment, two or three auxiliary tasks in the landmark rotation bounding box prediction task, the landmark semantic prediction task, and the entity-landmark contrastive learning task can be used to optimize the semantic grid features, and the visual representation is obtained.

[0077] Specifically, in the case where the plurality of auxiliary prediction tasks includes the landmark rotation bounding box prediction task, the landmark semantic prediction task, and the entity-landmark contrastive learning task, based on the plurality of auxiliary prediction tasks, display learning of the fine-grained alignment relationship between the entity and the landmark object is performed according to the semantic grid features, and the visual representation is obtained, including:

[0078] (1) Through the landmark rotation bounding box prediction task, landmark rotation bounding box information is obtained according to the multi-head cross attention mechanism and the semantic grid features.

[0079] In this embodiment, the landmark rotation bounding box prediction task aggregates the semantic grid features through the multi-head cross attention mechanism, and uses a feedforward neural network to predict the rotation bounding box parameters of the landmark, and finally optimizes it through a Smooth L1 loss function.

[0080] In this embodiment, in the landmark rotation bounding box prediction task, the entity embedding after aggregating the semantic grid features is obtained by the multi-head cross attention mechanism using the following formula:

[0081] ;

[0082] Wherein, represents the entity embedding after aggregating the semantic grid features through the multi-head cross attention mechanism; is a text entity feature; represents the feature fusion of and by the multi-head cross attention module; is a feedforward neural network, which is used to further map the aggregated features.

[0083] In this embodiment, the formula for predicting the rotated bounding box is as follows:

[0084] ;

[0085] wherein, denotes the parameters of the predicted landmark rotated bounding box, which is mapped from the features by the feedforward neural network and normalized to a reasonable range by using the Sigmoid activation function.

[0086] In this embodiment, the landmark rotated bounding box prediction task adopts a multi-head cross-attention mechanism to extract features from the semantic grid features and predict the landmark rotated bounding box The loss function corresponding to the landmark rotated bounding box prediction task is defined as:

[0087] ;

[0088] wherein, is the predicted bounding box, is the real bounding box. The loss function ensures that the model can learn the accurate geometric structure and rotation information of the landmark; the landmark rotated bounding box prediction task improves the detection and rotation adaptation ability of the model for complex geometric shape landmarks by integrating semantic information and visual features.

[0089] (2) The landmark semantic prediction task obtains the landmark language description information according to the autoregressive mechanism and the semantic grid features.

[0090] In this embodiment, the landmark semantic prediction task inputs the landmark features extracted from the semantic grid into the Transformer decoder to generate text descriptions related to the landmark, and optimizes the semantic prediction accuracy through the cross-entropy loss function.

[0091] In this embodiment, the landmark semantic prediction task generates the semantic description of the landmark through the following formula:

[0092] ;

[0093] wherein, denotes the predicted generated semantic description, is the feature extracted from the semantic grid by the rotated RoI-Align operation. The Transformer decoder generates text descriptions related to the landmark through the autoregressive mechanism.

[0094] The loss function formula corresponding to the landmark semantic prediction task is:

[0095] ;

[0096] wherein, a semantic description generated for the model, a semantic description for the real landmark. This task combines the regional features of the aerial image with the visual features of the landmark, and generates a semantically accurate language description through a Transformer decoder.

[0097] (3) Through the entity-landmark contrast learning task, the cross-modal features of entities and landmarks are compared and learned according to the landmark rotation bounding box information and the landmark language description information, and the visual representation is obtained.

[0098] In this embodiment, the entity-landmark contrast learning task optimizes the alignment relationship between the language entity and the landmark visual feature through the contrast learning method, calculates the contrast loss of the entity to the landmark and the landmark to the entity using a bidirectional contrast loss function, and improves the alignment accuracy of the model in the multi-modal feature space.

[0099] In this embodiment, the entity-landmark contrast learning task optimizes the cross-modal features of the language entity and the visual landmark through the contrast learning method, and the corresponding loss function is:

[0100] ;

[0101] ;

[0102] The above formula describes the bidirectional optimization process of contrast learning, wherein, and are the contrast losses from the entity to the landmark and from the landmark to the entity, respectively; is the text entity feature, is the landmark visual feature; is a temperature parameter that controls the smoothness of the probability distribution; is the matching label, indicating and whether they are a matching pair; the navigation model is trained through the entity-landmark contrast learning task, and the navigation model can achieve fine-grained entity-landmark alignment in the multi-modal feature space.

[0103] Step 130, taking the fine-grained aerial visual dialogue navigation dataset as the training sample, taking the visual representation as the input feature, and taking the comprehensive loss as the loss function, iteratively training the navigation model to obtain an aerial visual dialogue navigation model; wherein the comprehensive loss is determined based on the navigation loss function and the loss function corresponding to the multiple auxiliary prediction tasks; the fine-grained aerial visual dialogue navigation dataset is obtained by fine-grained expansion of the entity-landmark data in the unmanned aerial vehicle dialogue navigation AVDN dataset.

[0104] In this step, the comprehensive loss is designed by jointly optimizing the navigation task and the auxiliary task, which can improve the navigation ability and multi-modal alignment accuracy of the model. Specifically, the navigation task optimizes the navigation strategy of the model by minimizing the mean square error between the predicted action and the real action, and the comprehensive loss combines the losses of the navigation task and the alignment task to improve the performance of the model in complex scenes through multi-task learning.

[0105] In this embodiment, the entity-landmark data in the AVDN dataset is expanded as follows: more landmark types are added, the position information of the landmark is refined, more diverse dialogue scenes and navigation instructions are added, etc., to obtain a refined aerial visual dialogue navigation dataset, which improves the diversity and richness of the dataset.

[0106] In this embodiment, the navigation model includes a neural network architecture for processing visual and text inputs, which includes an encoder-decoder structure, wherein the encoder is used to process input features (images and texts), and the decoder is used to generate navigation instructions or predict the next action; in addition, the navigation model can help the model pay more attention to important visual and text information when generating navigation instructions by introducing an attention mechanism.

[0107] In this embodiment, the comprehensive loss is determined by the following formula:

[0108] ;

[0109] wherein, is the comprehensive loss, is the navigation loss function, is the loss function corresponding to each auxiliary prediction task;

[0110] is determined by the following formula:

[0111] ;

[0112] wherein, is the loss function corresponding to the landmark rotation bounding box prediction task, is the loss function corresponding to the landmark semantic prediction task, is the loss function corresponding to the entity-landmark contrastive learning task; and are weight parameters for balancing the loss weights of each task.

[0113] In this embodiment, the navigation loss function is determined by the following formula:

[0114] ;

[0115] wherein, representing the navigation action predicted by the model, representing the true label of the navigation, For a time step, the navigation loss optimizes the navigation policy of the model by minimizing the mean squared error between the predicted action and the true action.

[0116] In this embodiment, before training the navigation model with the samples, the samples can be preprocessed through operations including but not limited to data cleaning, normalization, enhancement, etc.; in the iteration process, the model can be iteratively trained using an optimization algorithm such as Adam, SGD, etc.; in each iteration, the comprehensive loss is calculated and the model parameters are updated; during the training process, the performance of the model is periodically evaluated, and the validation set or test set can be used to check the generalization ability of the model; after the training is completed, the hyperparameters of the model (such as learning rate, batch size, etc.) can be adjusted according to the evaluation results to optimize the performance.

[0117] The navigation model training method for fine alignment learning provided by the embodiment of the application, by weighting and fusing the visual features, semantic features and spatial features extracted from the aerial images of the unmanned aerial vehicle, obtaining semantic grid features, and using multiple auxiliary prediction tasks to display and learn the fine alignment relationship between entities and landmark objects according to the semantic grid features, obtaining visual representations, finally taking the fine aerial visual dialogue navigation dataset as the training sample, taking the visual representations as the input features, and taking the comprehensive loss as the loss function to iteratively train the navigation model, obtaining the aerial visual dialogue navigation model, through comprehensive fusion of multi-modal features (including visual, semantic and spatial information), the navigation accuracy, alignment capability and task execution efficiency of the unmanned aerial vehicle in complex scenes are improved.

[0118] In some embodiments, the fine aerial visual dialogue navigation dataset is obtained by the following steps:

[0119] (1) Extracting language entities from the navigation dialogue data of the AVDN dataset; the language entities include multiple landmark objects, landmark orientation descriptions and semantic relationships.

[0120] In this embodiment, language entities are extracted from the navigation dialogue by a natural language processing model to identify the names, orientation descriptions and semantic relationships of target landmarks, so as to enhance the positioning effect of language description on landmarks.

[0121] In this embodiment, the extracted language entities include the names, attributes and relative position descriptions of target landmarks to accurately reflect the semantic information in the navigation dialogue and ensure that the navigation model can effectively capture the context relationship between landmark information and unmanned aerial vehicle operation. The language entities need to be structurally analyzed by dependency parsing technology to enhance the processing capability of complex dialogue structures.

[0122] In this embodiment, before extracting language entities, the navigation dialogue data can be cleaned by removing irrelevant characters in the dialogue text, such as punctuation marks, special symbols, etc., and the dialogue text is segmented into words or phrases using a word segmentation tool (such as jieba, NLTK, etc.), and then part-of-speech tagging is performed to identify key information such as nouns, verbs, and directional words.

[0123] In this embodiment, the natural language processing model is used to identify named entities in the navigation dialogue data, particularly names related to landmarks, such as landmark colors, quantities, and shapes, etc.

[0124] In this embodiment, the landmark orientation description can be determined by analyzing the semantic relationship between the orientation words and the landmark names.

[0125] In this embodiment, the semantic relationship can be obtained by the following steps: using a dependency syntax analysis tool to analyze the syntax structure of the dialogue text and identify the dependency relationship between the landmark name and other words; based on the results of the dependency syntax analysis, extracting the semantic relationship between the landmark name and the orientation description, action instruction, etc.; for example, in "land in front of the red landmark", "land" is the action instruction, "red landmark" is the landmark name, and "front" is the orientation description.

[0126] (2) Obtain the landmark mask from the AVDN dataset and extract the rotation bounding box information of the landmark mask, including the landmark center point, width, height, and rotation angle.

[0127] In this embodiment, the semantic segmentation model can be used to generate the landmark mask and extract the rotation bounding box information of the landmark, including the center point, width, height, and rotation angle, to clearly define the geometric characteristics of the landmark.

[0128] In this embodiment, the semantic segmentation model includes but is not limited to U-Net, DeepLabV3+, and FPN, etc.

[0129] In this embodiment, the trained semantic segmentation model is used to predict the images in the VDN dataset to generate semantic segmentation masks of landmarks; in the semantic segmentation masks, the pixel values of the landmark regions are usually 1 (or a certain specific value), and the pixel values of other regions are 0; preprocessing such as denoising, morphological operations (dilation, erosion, opening operation, closing operation) is performed on the generated masks to optimize the extraction effect of the bounding box; then the contour detection function in the image processing library (such as OpenCV) is used to detect the contour of the landmark on the preprocessed mask, and the minimum enclosing rotating rectangle algorithm (such as the minAreaRect function of OpenCV) is used to fit the rotating bounding box of the detected contour to obtain the center point coordinates, width, height and rotation angle of the bounding box; finally, the language entity and the landmark mask are preliminarily aligned, and the misaligned entity-landmark pairs are corrected by manual verification to obtain refined aerial visual dialogue navigation data.

[0130] The navigation model training method for refined alignment learning provided by the embodiment of the application can obtain refined aerial visual dialogue navigation data containing detailed semantic annotations, landmark geometric features and multi-modal information, and can correct errors in entity-landmark alignment through an artificial verification mechanism, thereby ensuring the diversity and high quality of the data and providing reliable data support for training and evaluating the navigation model.

[0131] In some embodiments, after obtaining the aerial visual dialogue navigation model, the method further includes: fusing and deciding the navigation result output by the aerial visual dialogue navigation model based on a time sequence transformation-based navigation strategy model to obtain a target navigation strategy; wherein the time sequence transformation-based navigation strategy model is used to determine a semantic grid feature, existing dialogue data, visual features, trajectory data and a predicted flight trajectory of a UAV.

[0132] In this embodiment, to cope with multi-modal navigation tasks in complex scenes, the time sequence transformation-based navigation strategy model combines dialogue history , visual history , trajectory history and semantic grid features to generate three-dimensional navigation actions of the UAV based on multi-modal feature fusion, so that the trained navigation model can efficiently predict a navigation path in a complex scene and has strong robustness and real-time performance.

[0133] In this embodiment, the time sequence transformation-based navigation strategy model combines dialogue history Visual history Trajectory History Multimodal historical information and semantic grid features Navigation actions are generated through a feedforward neural network. The calculation formula is:

[0134] ;

[0135] in, It represents the three-dimensional spatial displacement of the UAV; by fusing multimodal features, the model can generate high-precision navigation paths in complex scenarios; the navigation model has the ability to efficiently handle multimodal navigation tasks in complex scenarios, and has both high precision and real-time inference performance, making it suitable for a variety of aerial navigation applications.

[0136] The refined alignment learning navigation model training method provided in this invention fuses and makes decisions on the navigation results output by the aerial visual dialogue navigation model through a navigation strategy model based on temporal transformation, thereby obtaining a target navigation strategy. It achieves efficient multimodal information fusion and feature extraction. Compared with traditional navigation strategy models, the transformer structure better captures long-range dependencies and local contextual information, thereby generating more accurate and robust navigation paths in complex scenarios.

[0137] Figure 2 This is the second flowchart illustrating the refined alignment learning navigation model training method provided by this invention. Figure 2 In the illustrated embodiment, an ORCNN network is used to extract aerial images from drones. Extracting landmark object codes (corresponding semantic features) from [the source], by [from] The location information of the drone extracted is used to obtain polar coordinate embedding (corresponding spatial features), and then... Visual encoding is performed to obtain latent features (corresponding to visual features). Then, the three types of features mentioned above are weighted and calculated, and then combined to obtain semantic grid features. Then, by utilizing the landmark rotation bounding box prediction task, the landmark semantic prediction task, and the entity-landmark comparison learning task, based on... The system displays learned entity-landmark alignment relationships, including using a pre-set FG-AVDN dataset as training samples (e.g., input "flying over a C-shaped wheat building and the southernmost white airplane to reach the destination"). The jointly constructed comprehensive loss is used to iteratively train the navigation model. The navigation prediction results and the text encoding results corresponding to the input text are jointly input into the navigation strategy model based on time-series transformation to complete the fusion and decision of multimodal features, and finally generate the navigation strategy, that is, to obtain the UAV action prediction results and progress prediction results.

[0138] The navigation method provided by the present application is described below. The navigation method described below can be correspondingly referred to the fine alignment learning navigation model training method described above.

[0139] Figure 3 is a flowchart of the navigation method provided by the present application, as Figure 3 indicated, the navigation method comprises the following steps:

[0140] Step 310: obtaining a UAV aerial image to be processed.

[0141] In this step, the UAV aerial image to be processed can be an image obtained from a fine aerial visual dialogue navigation data set or a UAV aerial image database, or can be a real-time image.

[0142] In this embodiment, the UAV aerial image to be processed includes airport, UAV, building and landmark, etc.

[0143] Step 320: processing the UAV aerial image to be processed based on an aerial visual dialogue navigation model to obtain a navigation result; the aerial visual dialogue navigation model is obtained by training the navigation model training method of fine alignment learning.

[0144] In this step, the aerial visual dialogue navigation model is obtained by training based on the following steps:

[0145] (1) extracting visual features, semantic features and spatial features from the UAV aerial image, and performing weighted fusion on the visual features, semantic features and spatial features to obtain semantic grid features; wherein the semantic features are used to represent the class information and geometric information of the landmark objects in the UAV aerial image; the spatial features are used to represent the relative direction and relative distance between the UAV and the landmark objects;

[0146] (2) based on a plurality of auxiliary prediction tasks, the fine alignment relationship between entities and landmark objects is displayed and learned based on the semantic grid features to obtain visual representation; wherein the plurality of auxiliary prediction tasks include at least two of landmark rotation bounding box prediction task, landmark semantic prediction task and entity-landmark contrast learning task;

[0147] (3) taking the fine aerial visual dialogue navigation data set as a training sample, taking the visual representation as an input feature, and taking a comprehensive loss as a loss function to iteratively train the navigation model to obtain an aerial visual dialogue navigation model; wherein the comprehensive loss is determined based on a navigation loss function and a loss function corresponding to the plurality of auxiliary prediction tasks; the fine aerial visual dialogue navigation data set is obtained by fine-grained expansion of entity-landmark data in the UAV dialogue navigation AVDN data set.

[0148] It should be noted that the implementation of each training step corresponds to steps 110-130 described above, and this embodiment will not be described again.

[0149] Step 330, the navigation result is fused and decided based on the navigation strategy model based on time sequence transformation, and the navigation strategy is obtained; wherein the navigation strategy model based on time sequence transformation is used to predict the navigation trajectory of the unmanned aerial vehicle according to the semantic grid feature, the existing dialogue data, the visual feature and the trajectory data.

[0150] In this embodiment, the navigation strategy model based on time sequence transformation is used to combine the dialogue history , the visual history , the trajectory history and the semantic grid feature , and generate the three-dimensional navigation action of the unmanned aerial vehicle based on the multi-modal feature fusion, so that the trained navigation model can efficiently predict the navigation path in a complex scene; the calculation formula of the predicted path of the unmanned aerial vehicle is:

[0151] ;

[0152] Wherein, represents the three-dimensional space displacement of the unmanned aerial vehicle; by fusing multi-modal features, the model can generate a high-precision navigation path in a complex scene; the navigation model has the ability to efficiently handle multi-modal navigation tasks in a complex scene, and has high precision and real-time inference performance, and is suitable for various air navigation applications.

[0153] In this embodiment, the to-be-processed unmanned aerial vehicle aerial image includes airport, unmanned aerial vehicle, building and landmark and the like, the to-be-processed unmanned aerial vehicle aerial image is processed by the air visual dialogue navigation model, the unmanned aerial vehicle navigation strategy is obtained, and the three-dimensional navigation action of the unmanned aerial vehicle is generated based on the multi-modal feature fusion by combining the dialogue history, the visual history, the trajectory history and the semantic grid feature corresponding to the unmanned aerial vehicle navigation path A, and a more accurate unmanned aerial vehicle navigation strategy is generated.

[0154] The navigation method provided by the embodiment of the application processes the to-be-processed unmanned aerial vehicle aerial image based on the air visual dialogue navigation model, obtains the navigation result, and then uses the navigation strategy model based on time sequence transformation to fuse and decide the navigation result, so as to obtain the navigation strategy, improve the navigation accuracy and robustness of the unmanned aerial vehicle in a complex dynamic scene, and can be widely used in practical application scenarios such as urban logistics, agricultural monitoring, search and rescue and industrial inspection.

[0155] The navigation model training device for fine alignment learning provided by the present application is described below, and the navigation model training device for fine alignment learning described below can be correspondingly referred to the navigation model training method for fine alignment learning described above.

[0156] Figure 4 Fig. 1 is a structural schematic diagram of the navigation model training device for fine alignment learning provided by the present application, as shown in the figure, the navigation model training device for fine alignment learning comprises a feature fusion module 410, an alignment learning module 420 and a training module 430. Figure 4

[0157] The feature fusion module 410 is configured to extract visual features, semantic features and spatial features from the UAV aerial image, and to perform weighted fusion on the visual features, the semantic features and the spatial features to obtain semantic grid features; wherein the semantic features are configured to represent the class information and the geometric information of the landmark object in the UAV aerial image; and the spatial features are configured to represent the relative direction and the relative distance between the UAV and the landmark object.

[0158] The alignment learning module 420 is configured to perform explicit learning on the fine alignment relationship between the entity and the landmark object based on a plurality of auxiliary prediction tasks according to the semantic grid features to obtain visual representations; wherein the plurality of auxiliary prediction tasks comprises at least two of the landmark rotation bounding box prediction task, the landmark semantic prediction task and the entity-landmark contrastive learning task.

[0159] The training module 430 is configured to take the fine aerial visual dialogue navigation data set as the training sample, take the visual representations as the input features, and take the comprehensive loss as the loss function to iteratively train the navigation model, so as to obtain an aerial visual dialogue navigation model; wherein the comprehensive loss is determined based on the navigation loss function and the loss functions corresponding to the plurality of auxiliary prediction tasks; and the fine aerial visual dialogue navigation data set is obtained by performing fine-grained expansion on the entity-landmark data in the UAV dialogue navigation AVDN data set.

[0160] The navigation model training device for fine alignment learning provided by the embodiment of the present application, by performing weighted fusion on the visual features, the semantic features and the spatial features extracted from the UAV aerial image to obtain semantic grid features, and by performing explicit learning on the fine alignment relationship between the entity and the landmark object based on a plurality of auxiliary prediction tasks according to the semantic grid features to obtain visual representations, finally taking the fine aerial visual dialogue navigation data set as the training sample, taking the visual representations as the input features, and taking the comprehensive loss as the loss function to iteratively train the navigation model, so as to obtain an aerial visual dialogue navigation model, through comprehensive fusion of multi-modal features (including visual, semantic and spatial information), the navigation precision, the alignment capability and the task execution efficiency of the UAV in a complex scene are improved.

[0161] ​The navigation device provided by the present application is described below, and the navigation device described below can be correspondingly referred to the navigation method described above.

[0162] Figure 5 The navigation device provided by the present application is described below, and the navigation device described below can be correspondingly referred to the navigation method described above. Figure 5 As shown in the figure, the navigation device comprises an image acquisition module 510, a navigation prediction module 520 and a navigation decision module 530.

[0163] The image acquisition module 510 is configured to acquire a UAV aerial image to be processed.

[0164] The navigation prediction module 520 is configured to process the UAV aerial image to be processed based on an aerial visual dialogue navigation model to obtain a navigation result; the aerial visual dialogue navigation model is obtained by training an aerial visual dialogue navigation model through a fine alignment learning navigation model training method.

[0165] The navigation decision module 530 is configured to fuse and decide the navigation result based on a navigation strategy model based on time sequence transformation to obtain a navigation strategy; wherein the navigation strategy model based on time sequence transformation is used to determine a semantic grid feature, existing dialogue data, visual features and trajectory data and a predicted UAV navigation trajectory.

[0166] The navigation device provided by the present application is described below, and the navigation device described below can be correspondingly referred to the navigation method described above.

[0167] Figure 6 The navigation device provided by the present application is described below, and the navigation device described below can be correspondingly referred to the navigation method described above. Figure 6As shown, the electronic device can include a processor 610, a communications interface 620, a memory 630, and a communications bus 640, wherein the processor 610, the communications interface 620, and the memory 630 communicate with each other through the communications bus 640. The processor 610 can invoke a logical instruction in the memory 630 to execute the navigation model training method of fine-grained alignment learning, which includes extracting visual features, semantic features, and spatial features from a UAV aerial image, and performing weighted fusion on the visual features, the semantic features, and the spatial features to obtain semantic grid features; wherein the semantic features are used to represent the class information and geometric information of landmark objects in the UAV aerial image; the spatial features are used to represent the relative direction and relative distance between the UAV and the landmark objects; based on a plurality of auxiliary prediction tasks, the semantic grid features are used to display learning of the fine-grained alignment relationship between entities and landmark objects to obtain visual representations; wherein the plurality of auxiliary prediction tasks include at least two of landmark rotation bounding box prediction task, landmark semantic prediction task, and entity-landmark contrast learning task; taking a fine-grained aerial visual dialogue navigation dataset as a training sample, taking the visual representation as an input feature, and taking a comprehensive loss as a loss function, the navigation model is iteratively trained to obtain an aerial visual dialogue navigation model; wherein the comprehensive loss is determined based on a navigation loss function and a loss function corresponding to the plurality of auxiliary prediction tasks; the fine-grained aerial visual dialogue navigation dataset is obtained by fine-grained expansion of entity-landmark data in a UAV dialogue navigation AVDN dataset.

[0168] Or execute a navigation method, the navigation method includes: obtaining a UAV aerial image to be processed; processing the UAV aerial image to be processed based on an aerial visual dialogue navigation model to obtain a navigation result; the aerial visual dialogue navigation model is trained by a navigation model training method of fine-grained alignment learning; a navigation strategy is obtained by fusing and deciding the navigation result based on a navigation strategy model based on time series transformation; wherein the navigation strategy model based on time series transformation is used to determine the navigation trajectory of the predicted UAV according to the semantic grid features, the existing dialogue data, the visual features, and the trajectory data.

[0169] Moreover, the logic instructions in the memory 630 described above can be implemented in the form of software functional units and sold or used as independent products, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0170] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor to enable a computer to execute the navigation model training method of fine alignment learning provided by the above-mentioned methods. The method comprises: extracting visual features, semantic features and spatial features from a UAV aerial image, and performing weighted fusion on the visual features, semantic features and spatial features to obtain semantic grid features; wherein the semantic features are used to represent the class information and geometric information of landmark objects in the UAV aerial image; the spatial features are used to represent the relative direction and relative distance between the UAV and the landmark objects; based on a plurality of auxiliary prediction tasks, the fine alignment relationship between entities and landmark objects is displayed and learned based on the semantic grid features to obtain visual representations; wherein the plurality of auxiliary prediction tasks include at least two of landmark rotation bounding box prediction tasks, landmark semantic prediction tasks and entity-landmark contrast learning tasks; taking a fine aerial visual dialogue navigation data set as a training sample, taking the visual representations as input features, and taking a comprehensive loss as a loss function, the navigation model is iteratively trained to obtain an aerial visual dialogue navigation model; wherein the comprehensive loss is determined based on a navigation loss function and a loss function corresponding to the plurality of auxiliary prediction tasks; the fine aerial visual dialogue navigation data set is obtained by fine-grained expansion of entity-landmark data in a UAV dialogue navigation AVDN data set.

[0171] Or a navigation method is performed, the navigation method comprising: obtaining a to-be-processed aerial image of a UAV; processing the to-be-processed aerial image of the UAV based on an aerial visual dialogue navigation model to obtain a navigation result; the aerial visual dialogue navigation model is trained by a navigation model training method of fine-grained alignment learning; a navigation strategy model based on time sequence transformation is used for fusing and deciding the navigation result to obtain a navigation strategy; and the navigation strategy model based on time sequence transformation is used for obtaining a navigation trajectory of the UAV according to semantic grid features, existing dialogue data, visual features, and trajectory data.

[0172] In another aspect, the application also provides a non-transitory computer-readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the navigation model training method of fine-grained alignment learning provided by the above method, and the method comprises: extracting visual features, semantic features, and spatial features from an aerial image of a UAV, and performing weighted fusion on the visual features, the semantic features, and the spatial features to obtain semantic grid features; the semantic features are used for representing class information and geometric information of landmark objects in the aerial image of the UAV; the spatial features are used for representing relative directions and relative distances between the UAV and the landmark objects; performing display learning on fine-grained alignment relationships between entities and landmark objects based on the semantic grid features according to a plurality of auxiliary prediction tasks to obtain visual representations; the plurality of auxiliary prediction tasks comprise at least two of a landmark rotation bounding box prediction task, a landmark semantic prediction task, and an entity-landmark contrast learning task; using a fine-grained aerial visual dialogue navigation dataset as a training sample, using the visual representations as input features, and using a comprehensive loss as a loss function to iteratively train a navigation model to obtain an aerial visual dialogue navigation model; the comprehensive loss is determined based on a navigation loss function and loss functions corresponding to the plurality of auxiliary prediction tasks; and the fine-grained aerial visual dialogue navigation dataset is obtained by performing fine-grained expansion on entity-landmark data in an aerial visual dialogue navigation (AVDN) dataset.

[0173] Or a navigation method is performed, the navigation method comprising: obtaining a to-be-processed aerial image of a UAV; processing the to-be-processed aerial image of the UAV based on an aerial visual dialogue navigation model to obtain a navigation result; the aerial visual dialogue navigation model is trained by a navigation model training method of fine-grained alignment learning; a navigation strategy model based on time sequence transformation is used for fusing and deciding the navigation result to obtain a navigation strategy; and the navigation strategy model based on time sequence transformation is used for obtaining a navigation trajectory of the UAV according to semantic grid features, existing dialogue data, visual features, and trajectory data.

[0174] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected to achieve the purposes of the embodiments according to actual needs. Those skilled in the art can understand and implement without creative labor.

[0175] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software and the necessary general hardware platform, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.

[0176] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A navigation model training method of fine-grained alignment learning, characterized by, Comprise: extracting visual features, semantic features and spatial features from the aerial image of the unmanned aerial vehicle, and performing weighted fusion on the visual features, the semantic features and the spatial features to obtain semantic grid features; wherein the semantic features are used to represent the class information and geometric information of landmark objects in the aerial image of the unmanned aerial vehicle; the spatial features are used to represent the relative direction and relative distance between the unmanned aerial vehicle and the landmark objects; based on a plurality of auxiliary prediction tasks, the semantic grid features are used to display learning of the fine alignment relationship between entities and landmark objects, and visual representation is obtained; wherein the plurality of auxiliary prediction tasks include at least two of landmark rotation bounding box prediction task, landmark semantic prediction task and entity-landmark contrast learning task; taking the fine aerial visual dialogue navigation dataset as the training sample, taking the visual representation as the input feature, and taking the comprehensive loss as the loss function, the navigation model is iteratively trained to obtain an aerial visual dialogue navigation model; wherein the comprehensive loss is determined based on the navigation loss function and the loss function corresponding to the plurality of auxiliary prediction tasks; the fine aerial visual dialogue navigation dataset is obtained by fine-grained expansion of entity-landmark data in the unmanned aerial vehicle dialogue navigation AVDN dataset; the fine aerial visual dialogue navigation dataset is obtained by the following steps: extracting language entities from the navigation dialogue data of the AVDN dataset; the language entities include a plurality of landmark objects, landmark orientation descriptions and semantic relationships; obtaining landmark masks from the AVDN dataset and extracting rotation bounding box information of the landmark masks, the rotation bounding box information including landmark center points, width, height and rotation angle; preliminarily aligning the language entities and the landmark masks, and correcting the mismatched entity-landmark pairs through artificial verification to obtain the fine aerial visual dialogue navigation data; the based on a plurality of auxiliary prediction tasks, the semantic grid features are used to display learning of the fine alignment relationship between entities and landmark objects, and visual representation is obtained; wherein the plurality of auxiliary prediction tasks include at least two of landmark rotation bounding box prediction task, landmark semantic prediction task and entity-landmark contrast learning task; the landmark rotation bounding box information is obtained through the landmark rotation bounding box prediction task based on the multi-head cross attention mechanism and the semantic grid features. 2.The fine-alignment learning-based navigation model training method of claim 1, wherein, The plurality of auxiliary prediction tasks include the landmark rotation bounding box prediction task, the landmark semantic prediction task and the entity-landmark contrast learning task; the based on a plurality of auxiliary prediction tasks, the semantic grid features are used to display learning of the fine alignment relationship between entities and landmark objects, and visual representation is obtained; wherein the plurality of auxiliary prediction tasks include at least two of landmark rotation bounding box prediction task, landmark semantic prediction task and entity-landmark contrast learning task; the landmark language description information is obtained through the landmark semantic prediction task based on the self-recurrent mechanism and the semantic grid features; the cross-modal features of entities and landmarks are compared and learned through the entity-landmark contrast learning task based on the landmark rotation bounding box information and the landmark language description information to obtain the visual representation. 3.The fine-tuned alignment learning-based navigation model training method of claim 1, wherein, The comprehensive loss is determined by the following formula: wherein, is a comprehensive loss, is the navigation loss function, is a loss function corresponding to the plurality of auxiliary prediction tasks; is represented by the following formula: wherein, is a loss function corresponding to the landmark rotation bounding box prediction task, is a loss function corresponding to the landmark semantic prediction task, is a loss function corresponding to the entity-landmark contrastive learning task; and κ1 and κ2 are weight parameters for balancing the loss weights of each task.

4. The refined alignment learning-based navigation model training method of claim 1, wherein, After obtaining the aerial visual dialogue navigation model, the method further comprises: The navigation strategy model based on time sequence transformation fuses and decides the navigation result output by the aerial visual dialogue navigation model, to obtain a target navigation strategy; wherein the navigation strategy model based on time sequence transformation is used to determine the navigation trajectory of the predicted unmanned aerial vehicle according to the semantic grid feature, the existing dialogue data, the visual feature, the trajectory data, and the navigation trajectory of the predicted unmanned aerial vehicle.

5. A navigation method characterized by, The method comprises: acquiring an aerial image of an unmanned aerial vehicle to be processed; processing the aerial image of the unmanned aerial vehicle to be processed based on an aerial visual dialogue navigation model to obtain a navigation result; the aerial visual dialogue navigation model is trained by the navigation model training method of refined alignment learning according to any one of claims 1-4; a navigation strategy model based on time sequence transformation fuses and decides the navigation result, to obtain a navigation strategy; wherein the navigation strategy model based on time sequence transformation is used to determine the navigation trajectory of the predicted unmanned aerial vehicle according to the semantic grid feature, the existing dialogue data, the visual feature, the trajectory data, and the navigation trajectory of the predicted unmanned aerial vehicle.

6. A navigation model training apparatus for refined alignment learning, which applies the navigation model training method for refined alignment learning according to claim 1, characterized by The method comprises: a feature fusion module is configured to extract visual features, semantic features, and spatial features from the aerial image of the unmanned aerial vehicle, and to perform weighted fusion on the visual features, the semantic features, and the spatial features to obtain a semantic grid feature; wherein the semantic features are used to represent the class information and geometric information of landmark objects in the aerial image of the unmanned aerial vehicle; and the spatial features are used to represent the relative direction and relative distance between the unmanned aerial vehicle and the landmark objects; an alignment learning module is configured to perform explicit learning on the refined alignment relationship between entities and landmark objects based on the semantic grid feature according to a plurality of auxiliary prediction tasks, to obtain visual representations; wherein the plurality of auxiliary prediction tasks include at least two of a landmark rotation bounding box prediction task, a landmark semantic prediction task, and an entity-landmark contrast learning task; a training module is configured to use a refined aerial visual dialogue navigation dataset as a training sample, use the visual representations as input features, and use a comprehensive loss as a loss function to iteratively train a navigation model, to obtain an aerial visual dialogue navigation model; wherein the comprehensive loss is determined based on a navigation loss function and loss functions corresponding to the plurality of auxiliary prediction tasks; and the refined aerial visual dialogue navigation dataset is obtained by performing fine-grained expansion on entity-landmark data in an aerial visual dialogue navigation (AVDN) dataset.

7. A navigation device characterized by The method comprises: an image acquisition module is configured to acquire an aerial image of an unmanned aerial vehicle to be processed; a navigation prediction module is configured to process the aerial image of the unmanned aerial vehicle to be processed based on an aerial visual dialogue navigation model, to obtain a navigation result; the aerial visual dialogue navigation model is trained by the navigation model training method of refined alignment learning according to any one of claims 1-4; a navigation decision module is configured to fuse and decide the navigation result based on a navigation strategy model based on time sequence transformation, to obtain a navigation strategy; wherein the navigation strategy model based on time sequence transformation is used to determine the navigation trajectory of the predicted unmanned aerial vehicle according to the semantic grid feature, the existing dialogue data, the visual feature, the trajectory data, and the navigation trajectory of the predicted unmanned aerial vehicle.

8. An electronic device comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, The processor executes the computer program to implement the method according to any one of claims 1-5. 9.A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program, which is executed by a processor, implements the method as claimed in any of claims 1 to 5.

Citation Information

Patent Citations

  • Navigation device

    CH659320A5

  • Method for determining pointing deviation of remote sensing satellite on geostationary orbit

    CN108364279A