Navigation model training method, navigation method and device for refined alignment learning
Through the navigation model training method of refined alignment learning, the visual, semantic and spatial features in the aerial images of drone are weighted and fusion, and the display learning of multiple auxiliary prediction tasks is carried out, which solves the problem of inaccurate prediction results in complex scenarios in the prior art, and achieves more efficient and accurate navigation path planning.
Patent Information
- Application Number
- CN202510043915.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-10
AI Technical Summary
Existing UAV navigation methods rely on coarse-grained global feature matching, making it difficult to accurately understand the semantic description and positional relationship of multiple landmarks in complex scenarios, resulting in inaccurate navigation prediction results.
The navigation model training method of refined alignment learning is adopted. By extracting visual features, semantic features and spatial features from drone aerial images for weighted fusion, semantic grid features are obtained, and a number of auxiliary prediction tasks are used to display and learn the fine alignment relationship between entities and landmark objects, and visual representation is obtained.
It improves the navigation accuracy and alignment capabilities of the drone in complex scenarios, enhances the task execution efficiency, and ensures the accuracy of navigation path planning.
Smart Images

Figure CN120014488A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a navigation model training method, a navigation method and a device for refined alignment learning. Background Art
[0002] In recent years, multimodal learning has gradually become an important research direction in the field of artificial intelligence. Its core goal is to enhance the model's understanding and reasoning capabilities in complex tasks by integrating visual and language information. In this context, the Aerial Vision-Dialog Navigation (AVDN) task, as a complex task that combines visual perception, language understanding and navigation planning, has received widespread attention.
[0003] In the related technologies, most of the existing UAV navigation methods rely on coarse-grained global feature matching, lacking fine-grained capture of visual features and language semantics, especially in complex scenes, the ability to understand the semantic descriptions and positional relationships of multiple landmarks is weak, resulting in inaccurate UAV navigation prediction results; on the other hand, when processing complex instructions, it is difficult for UAVs to simultaneously capture the combined information such as entity attributes in the language, relationships between landmarks, and priorities, and due to the lack of refined modeling of landmark semantics, especially in dynamic navigation tasks, it is difficult to achieve real-time interaction and decision-making of multimodal information, resulting in inaccurate navigation path planning. Summary of the invention
[0004] The present invention provides a navigation model training method, a navigation method and a device for refined alignment learning, so as to solve the defect that the drone navigation method in the prior art relies on coarse-grained global feature matching, has a weak ability to understand the semantic description and positional relationship of multiple landmarks in complex scenes, resulting in inaccurate drone navigation prediction results, and the prior art is difficult to simultaneously capture the entity attributes in the language, the relationship between landmarks, and the priority and other combined information, resulting in inaccurate navigation path planning, thereby improving the efficiency and accuracy of drone navigation planning.
[0005] The present invention provides a navigation model training method for refined alignment learning, comprising: Extracting visual features, semantic features and spatial features from the drone aerial image, and weighted fusion of the visual features, the semantic features and the spatial features to obtain semantic mesh features; wherein the semantic features are used to represent the category information and geometric information of the landmark object in the drone aerial image; and the spatial features are used to represent the relative direction and relative distance between the drone and the landmark object; Based on multiple auxiliary prediction tasks, the refined alignment relationship between the entity and the landmark object is displayed and learned according to the semantic grid features to obtain a visual representation; wherein the multiple auxiliary prediction tasks include at least two of a landmark rotation bounding box prediction task, a landmark semantic prediction task, and an entity-landmark comparison learning task; The refined aerial visual dialogue navigation dataset is used as a training sample, the visual representation is used as an input feature, and the navigation model is iteratively trained with a comprehensive loss as a loss function to obtain an aerial visual dialogue navigation model; wherein the comprehensive loss is determined based on the navigation loss function and the loss functions corresponding to the multiple auxiliary prediction tasks; the refined aerial visual dialogue navigation dataset is obtained by fine-grained expansion of the entity-landmark data in the UAV conversational navigation AVDN dataset.
[0006] According to a navigation model training method for refined alignment learning provided by the present invention, the refined aerial visual dialogue navigation dataset is obtained by the following steps: Extracting language entities from navigation dialogue data of the AVDN dataset; the language entities include a plurality of landmark objects, landmark location descriptions and semantic relationships; Obtain a landmark mask from the AVDN dataset, and extract the rotation bounding box information of the landmark mask, wherein the rotation bounding box information includes the landmark center point, width, height, and rotation angle; The language entity and the landmark mask are preliminarily aligned, and the mismatched entity-landmark pairs are corrected through manual verification to obtain the refined aerial visual dialogue navigation data.
[0007] According to a navigation model training method for refined alignment learning provided by the present invention, the multiple auxiliary prediction tasks include the landmark rotation bounding box prediction task, the landmark semantic prediction task and the entity-landmark comparison learning task; The display learning of the refined alignment relationship between the entity and the landmark object based on the semantic grid features based on multiple auxiliary prediction tasks to obtain the visual representation includes: Obtaining landmark rotation bounding box information through the landmark rotation bounding box prediction task according to the multi-head cross attention mechanism and the semantic grid feature; Acquire landmark language description information through the landmark semantic prediction task according to the autoregressive mechanism and the semantic grid features; The entity-landmark contrast learning task is used to contrast and learn the cross-modal features of the entity and the landmark according to the landmark rotation bounding box information and the landmark language description information to obtain the visual representation.
[0008] According to a navigation model training method for refined alignment learning provided by the present invention, the comprehensive loss is determined by the following formula: ; in, is the comprehensive loss, is the navigation loss function, is the loss function corresponding to the multiple auxiliary prediction tasks; It is expressed by the following formula: ; in, is the loss function corresponding to the landmark rotation bounding box prediction task, is the loss function corresponding to the landmark semantic prediction task, is the loss function corresponding to the entity-landmark comparison learning task; and is a weight parameter used to balance the loss weight of each task.
[0009] According to a navigation model training method for refined alignment learning provided by the present invention, after obtaining the aerial visual dialogue navigation model, the method further includes: The navigation strategy model based on time series transformation fuses and makes decisions on the navigation results output by the aerial visual dialogue navigation model to obtain a target navigation strategy; wherein the navigation strategy model based on time series transformation is used to predict the navigation trajectory of the UAV according to the semantic grid features, existing dialogue data, visual features and trajectory data.
[0010] The present invention also provides a navigation method, comprising: Obtain the drone aerial images to be processed; The unmanned aerial vehicle aerial image to be processed is processed based on the aerial visual dialogue navigation model to obtain a navigation result; the aerial visual dialogue navigation model is trained by the navigation model training method of the refined alignment learning; The navigation strategy model based on time series transformation fuses and makes decisions on the navigation results to obtain a navigation strategy; wherein the navigation strategy model based on time series transformation is used to predict the navigation trajectory of the drone according to semantic grid features, existing conversation data, visual features and trajectory data.
[0011] The present invention also provides a navigation model training device for refined alignment learning, comprising: A feature fusion module is used to extract visual features, semantic features and spatial features from the drone aerial images, and perform weighted fusion on the visual features, the semantic features and the spatial features to obtain semantic mesh features; wherein the semantic features are used to represent the category information and geometric information of the landmark objects in the drone aerial images; and the spatial features are used to represent the relative direction and relative distance between the drone and the landmark objects; An alignment learning module, configured to perform display learning on a refined alignment relationship between an entity and a landmark object according to the semantic grid features based on a plurality of auxiliary prediction tasks to obtain a visual representation; wherein the plurality of auxiliary prediction tasks include at least two of a landmark rotation bounding box prediction task, a landmark semantic prediction task, and an entity-landmark comparison learning task; A training module is used to iteratively train the navigation model using a refined aerial visual dialogue navigation dataset as a training sample, the visual representation as an input feature, and a comprehensive loss as a loss function to obtain an aerial visual dialogue navigation model; wherein the comprehensive loss is determined based on the navigation loss function and the loss functions corresponding to the multiple auxiliary prediction tasks; the refined aerial visual dialogue navigation dataset is obtained by fine-grained expansion of the entity-landmark data in the UAV conversational navigation AVDN dataset.
[0012] The present invention also provides a navigation device, comprising: An image acquisition module is used to acquire the drone aerial images to be processed; A navigation prediction module is used to process the unmanned aerial vehicle aerial image to be processed based on an aerial visual dialogue navigation model to obtain a navigation result; the aerial visual dialogue navigation model is trained by the navigation model training method of the refined alignment learning; A navigation decision module is used to fuse and make decisions on the navigation results based on a time-series transformation navigation strategy model to obtain a navigation strategy; wherein the time-series transformation-based navigation strategy model is used to predict the navigation trajectory of the drone based on semantic grid features, existing conversation data, visual features and trajectory data.
[0013] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it implements a navigation model training method or a navigation method for refined alignment learning as described in any one of the above.
[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a navigation model training method or a navigation method for refined alignment learning as described in any one of the above.
[0015] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements a navigation model training method or a navigation method for refined alignment learning as described in any one of the above.
[0016] The navigation model training method, navigation method and device for refined alignment learning provided by the present invention obtain semantic grid features by weighted fusion of visual features, semantic features and spatial features extracted from drone aerial images, and use multiple auxiliary prediction tasks to display and learn the refined alignment relationship between entities and landmark objects according to the semantic grid features to obtain visual representations. Finally, the navigation model is iteratively trained using a refined aerial visual dialogue navigation dataset as training samples, visual representations as input features, and comprehensive losses as loss functions to obtain an aerial visual dialogue navigation model. By comprehensively integrating multimodal features (including visual, semantic and spatial information), the navigation accuracy, alignment capability and task execution efficiency of drones in complex scenes are improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0018] Figure 1 It is one of the flow charts of the navigation model training method for refined alignment learning provided by the present invention.
[0019] Figure 2 This is the second flow chart of the navigation model training method for refined alignment learning provided by the present invention.
[0020] Figure 3 It is a flowchart diagram of the navigation method provided by the present invention.
[0021] Figure 4 It is a structural schematic diagram of a navigation model training device for refined alignment learning provided by the present invention.
[0022] Figure 5 It is a structural schematic diagram of the navigation device provided by the present invention.
[0023] Figure 6 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0024] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0025] Combine the following Figure 1-Figure 5 The present invention describes a navigation model training method, a navigation method and a device for refined alignment learning.
[0026] Figure 1 This is one of the flow charts of the navigation model training method for refined alignment learning provided by the present invention, such as Figure 1 As shown, the navigation model training method includes the following steps: Step 110: extract visual features, semantic features and spatial features from the drone aerial images, and perform weighted fusion on the visual features, semantic features and spatial features to obtain semantic mesh features; wherein the semantic features are used to represent the category information and geometric information of the landmark objects in the drone aerial images; and the spatial features are used to represent the relative direction and relative distance between the drone and the landmark objects.
[0027] In this step, drone aerial images can be obtained from AVDN (Aerial Vision-and-Dialog Navigation) data, which contains the corresponding drone-human natural language dialogue.
[0028] In this step, in order to improve the visual perception ability of the navigation model, visual, semantic and spatial features can be extracted from the aerial images of the drone, and semantic grid features can be constructed; the specific process is as follows: (1) Extracting visual features from drone aerial images , as the preliminary meshing features of the scene.
[0029] In this embodiment, the drone aerial image can be encoded by a visual encoder to obtain corresponding visual features.
[0030] In this embodiment, the visual encoder includes but is not limited to a convolutional neural network (CNN) or a neural network model based on a self-attention mechanism.
[0031] In this embodiment, before extracting features from the drone aerial images, the drone aerial images may be preprocessed including denoising, enhancement, correction, etc., to improve image quality and reduce the difficulty of subsequent processing.
[0032] (2) Use the target detection model to detect drone aerial images, generate the category labels, boundary information and geometric features of landmark objects, and further generate the semantic features of the objects. , to capture the category and geometric information of landmarks.
[0033] In this embodiment, the target detection model includes but is not limited to the YOLO (You Only Look Once) series and the R-CNN (Regions with Convolutional Neural Networks) series.
[0034] In this embodiment, a drone aerial image dataset is collected and annotated. The dataset may include various landmark objects, such as buildings, bridges, sculptures, etc., and their category labels, bounding boxes, and geometric features (such as height, width, area, etc.) are annotated.
[0035] In this embodiment, the trained target detection model is used to perform target detection on the preprocessed image. The model will identify the landmark object and generate its category label and bounding box. Then, based on the target detection result, the geometric features of the landmark object are extracted, such as height, width, area, perimeter, etc. These features can be used for subsequent analysis and processing.
[0036] In this embodiment, data annotation may be performed using methods such as polygon annotation and bounding box annotation to ensure the accuracy and completeness of the annotation.
[0037] In this embodiment, the category label, boundary information and geometric features of the landmark object can be integrated to form a feature vector containing the category information and geometric information of the landmark object; then the feature vector is further processed using deep learning technology (convolutional neural network or recurrent neural network) to generate semantic features of the landmark object for subsequent tasks such as landmark recognition, classification, retrieval, etc.
[0038] (3) Combine the location information of the drone to generate polar coordinate spatial features between the landmark and the drone , by describing the relative directions and distances of landmarks to improve the model’s spatial perception.
[0039] In this embodiment, the location information of the drone, such as longitude, latitude and altitude (or height) information, can be obtained through the positioning device of the drone (such as a GPS system).
[0040] In this embodiment, the polar coordinate space feature is generated based on the position information of the drone. It represents the relative relationship between the landmark and the drone, and its formula is: ; in, is the relative angle between the landmark and the drone, is the normalized landmark distance; the polar coordinate space encoding can explicitly capture the geometric relationship of the navigation path and enhance the model's perception of the spatial layout by standardizing the spatial distance and direction between the UAV and the target landmark.
[0041] (4) Weighted fusion of visual features, semantic coding and spatial coding to generate the final semantic mesh features ; It integrates visual, semantic and spatial information and can effectively capture the spatial layout and semantic relationships of landmarks in the scene.
[0042] In this embodiment, the semantic grid features The generation of , semantic encoding and polar space encoding , and its specific fusion method is: ; in, Representation layer normalization operation, , , The semantic grid feature combines three types of features, namely, visual, semantic and spatial encoding, so that the drone can effectively capture the semantic information and spatial distribution of landmarks in complex scenes.
[0043] Step 120, based on multiple auxiliary prediction tasks, the refined alignment relationship between the entity and the landmark object is displayed and learned according to the semantic grid features to obtain a visual representation; wherein the multiple auxiliary prediction tasks include at least two of the landmark rotation bounding box prediction task, the landmark semantic prediction task and the entity-landmark comparison learning task.
[0044] In this step, the landmark rotation bounding box prediction task can capture the geometric structure and rotation information of the landmark, the landmark semantic prediction task can obtain the text description information related to the landmark, and the entity-landmark comparison learning task can improve the alignment accuracy of the navigation model in the multimodal feature space.
[0045] In this embodiment, in order to enhance the model's ability in multimodal feature alignment, the semantic grid features can be optimized through two or three auxiliary tasks among the landmark rotation bounding box prediction task, the landmark semantic prediction task, and the entity-landmark comparison learning task to obtain a visual representation.
[0046] Specifically, in the case of multiple auxiliary prediction tasks including landmark rotation bounding box prediction task, landmark semantic prediction task and entity-landmark comparison learning task; based on multiple auxiliary prediction tasks, the refined alignment relationship between entity and landmark objects is displayed and learned according to semantic grid features, and the visual representations obtained include: (1) The landmark rotation bounding box prediction task is used to obtain the landmark rotation bounding box information based on the multi-head cross-attention mechanism and semantic grid features.
[0047] In this embodiment, the landmark rotation bounding box prediction task aggregates semantic grid features through a multi-head cross-attention mechanism, and uses a feedforward neural network to predict the landmark's rotation bounding box parameters, which is finally optimized through the Smooth L1 loss function.
[0048] In this embodiment, in the landmark rotation bounding box prediction task, the entity embedding after aggregating semantic grid features is obtained by the multi-head cross attention mechanism using the following formula: ; in, Represents the aggregation of semantic grids through multi-head cross attention mechanism Entity embedding after features; is the text entity feature; Represents a multi-head cross attention module pair and Feature fusion of It is a feed-forward neural network, which is used to further map the aggregated features.
[0049] In this embodiment, the formula for rotating the bounding box prediction is as follows: ; in, Represents the parameters of the predicted landmark rotation bounding box, maps the features through a feed-forward neural network, and uses the Sigmoid activation function to normalize the parameters to a reasonable range.
[0050] In this embodiment, the landmark rotation bounding box prediction task uses a multi-head cross attention mechanism to extract features from semantic grid features and predict the landmark rotation bounding box. ; The loss function corresponding to the rotation bounding box prediction task is defined as: ; in, To predict the bounding box, is the true bounding box. The loss function ensures that the model can learn the precise geometric structure and rotation information of the landmark; the landmark rotation bounding box prediction task improves the model's detection and rotation adaptation capabilities for landmarks with complex geometric shapes by integrating semantic information and visual features.
[0051] (2) The landmark semantic prediction task is used to obtain the landmark language description information based on the autoregressive mechanism and semantic grid features.
[0052] In this embodiment, the landmark semantic prediction task inputs the landmark features extracted from the semantic grid into the Transformer decoder, generates a text description related to the landmark, and optimizes the semantic prediction accuracy through the cross entropy loss function.
[0053] In this embodiment, the landmark semantic prediction task generates the semantic description of the landmark through the following formula: ; in, represents the semantic description of the prediction generation, To rotate the RoI-Align operation from the semantic mesh The Transformer decoder generates text descriptions related to landmarks through an autoregressive mechanism.
[0054] The loss function formula corresponding to the landmark semantic prediction task is: ; in, The semantic description generated for the model, This task combines the regional features of aerial images with the visual features of landmarks and generates semantically accurate language descriptions through the Transformer decoder.
[0055] (3) Through the entity-landmark contrast learning task, the cross-modal features of entities and landmarks are contrastively learned based on the landmark rotation bounding box information and the landmark language description information to obtain a visual representation.
[0056] In this embodiment, the entity-landmark contrastive learning task optimizes the alignment relationship between language entities and landmark visual features through a contrastive learning method, and uses a bidirectional contrastive loss function to calculate the contrast loss from entity to landmark and from landmark to entity, thereby improving the alignment accuracy of the model in the multimodal feature space.
[0057] In this embodiment, the entity-landmark contrastive learning task optimizes the cross-modal features of language entities and visual landmarks through a contrastive learning method, and the corresponding loss function is: ; ; The above formula describes the bidirectional optimization process of contrastive learning, where and They are the contrastive losses from entity to landmark and from landmark to entity respectively; is the text entity feature, It is the visual feature of the landmark; Control the smoothness of the probability distribution for the temperature parameter; To match the annotation, and Whether it is a matching pair; the navigation model is trained by the body-landmark comparison learning task, and the navigation model can achieve refined entity-landmark alignment in the multimodal feature space.
[0058] Step 130, using the refined aerial visual dialogue navigation dataset as training samples, visual representation as input features, and comprehensive loss as loss function to iteratively train the navigation model to obtain an aerial visual dialogue navigation model; wherein the comprehensive loss is determined based on the navigation loss function and the loss functions corresponding to multiple auxiliary prediction tasks; the refined aerial visual dialogue navigation dataset is obtained by fine-grained expansion of the entity-landmark data in the UAV conversational navigation AVDN dataset.
[0059] In this step, by jointly optimizing the navigation task and the auxiliary task to design a comprehensive loss, the navigation ability and multimodal alignment accuracy of the model can be improved. Specifically, the navigation task optimizes the navigation strategy of the model by minimizing the mean square error between the predicted action and the real action. The comprehensive loss combines the loss of the navigation task with the alignment task, and improves the performance of the model in complex scenarios through multi-task learning.
[0060] In this embodiment, by expanding the entity-landmark data in the AVDN dataset as follows: adding more landmark types, refining the location information of the landmarks, adding more diverse dialogue scenarios and navigation instructions, etc., a refined aerial visual dialogue navigation dataset is obtained, which improves the diversity and richness of the dataset.
[0061] In this embodiment, the navigation model includes a neural network architecture for processing visual and text inputs, and the neural network architecture includes an encoder-decoder structure, wherein the encoder is used to process input features (images and text), and the decoder is used to generate navigation instructions or predict the next action; in addition, the navigation model also introduces an attention mechanism to help the model pay more attention to important visual and text information when generating navigation instructions.
[0062] In this example, the comprehensive loss is determined by the following formula: ; in, is the comprehensive loss, is the navigation loss function, Loss functions corresponding to multiple auxiliary prediction tasks; Determined by the following formula: ; in, The loss function corresponding to the landmark rotation bounding box prediction task, is the loss function corresponding to the landmark semantic prediction task, is the loss function corresponding to the entity-landmark comparison learning task; and is a weight parameter used to balance the loss weight of each task.
[0063] In this embodiment, the navigation loss function Determined by the following formula: ; in, represents the navigation action predicted by the model, Indicates the actual label of the navigation. The navigation loss optimizes the model’s navigation strategy by minimizing the mean square error between the predicted action and the true action.
[0064] In this embodiment, before using samples to train the navigation model, the samples can be preprocessed through operations including but not limited to data cleaning, normalization, enhancement, etc.; during the iteration process, the model can be iteratively trained using an optimization algorithm (such as Adam, SGD, etc.); in each iteration, the comprehensive loss is calculated and the model parameters are updated; during the training process, the performance of the model is regularly evaluated, and a validation set or a test set can be used to check the generalization ability of the model; after the training is completed, the hyperparameters of the model (such as learning rate, batch size, etc.) can be adjusted based on the evaluation results to optimize the performance.
[0065] The navigation model training method for refined alignment learning provided in an embodiment of the present invention obtains semantic grid features by weighted fusion of visual features, semantic features and spatial features extracted from aerial images of unmanned aerial vehicles, and uses multiple auxiliary prediction tasks to perform explicit learning on the refined alignment relationship between entities and landmark objects according to the semantic grid features to obtain visual representations. Finally, the navigation model is iteratively trained using a refined aerial visual dialogue navigation dataset as training samples, visual representations as input features, and comprehensive loss as the loss function to obtain an aerial visual dialogue navigation model. By comprehensively integrating multimodal features (including visual, semantic and spatial information), the navigation accuracy, alignment capability and task execution efficiency of unmanned aerial vehicles in complex scenes are improved.
[0066] In some embodiments, the refined aerial visual dialogue navigation dataset is obtained by the following steps: (1) Extract language entities from the navigation dialogue data of the AVDN dataset; language entities include multiple landmark objects, landmark location descriptions, and semantic relations.
[0067] In this embodiment, language entities are extracted from the navigation dialogue through a natural language processing model to identify the name, location description and semantic relationship of the target landmark, so as to enhance the positioning effect of the language description on the landmark.
[0068] In this embodiment, the extracted language entities include the name, attributes and relative position description of the target landmark to accurately reflect the semantic information in the navigation dialogue and ensure that the navigation model can effectively capture the contextual relationship between landmark information and drone operations. The language entities need to be structured and analyzed through dependency parsing technology to enhance the ability to handle complex dialogue structures.
[0069] In this embodiment, before extracting language entities, the navigation conversation data can be cleaned by removing irrelevant characters in the conversation text, such as punctuation marks, special symbols, etc., and using word segmentation tools (such as jieba, NLTK, etc.) to segment the conversation text into words or phrases, and then using part-of-speech tagging to identify key information such as nouns, verbs, and directional words.
[0070] In this embodiment, a natural language processing model is used to identify named entities in the navigation conversation data, especially names related to landmarks, such as landmark colors, quantities, and shapes.
[0071] In this embodiment, the landmark location description can be determined by analyzing the semantic relationship between the location word and the landmark name.
[0072] In this embodiment, the semantic relationship can be obtained through the following steps: using a dependency syntax analysis tool to parse the syntactic structure of the dialogue text and identify the dependency relationship between the landmark name and other vocabulary; based on the results of the dependency syntax analysis, extracting the semantic relationship between the landmark name and the direction description, action instructions, etc.; for example, in "land in front of the red landmark", "landing" is an action instruction, "red landmark" is the landmark name, and "in front" is a direction description.
[0073] (2) Obtain the landmark mask from the AVDN dataset and extract the rotated bounding box information of the landmark mask. The rotated bounding box information includes the landmark center point, width, height, and rotation angle.
[0074] In this embodiment, a semantic segmentation model may be used to generate a landmark mask, and to extract the rotation bounding box information of the landmark, including the center point, width, height, and rotation angle, so as to clarify the geometric characteristics of the landmark.
[0075] In this embodiment, the semantic segmentation model includes but is not limited to U-Net, DeepLabV3+, and FPN.
[0076] In this embodiment, the trained semantic segmentation model is used to predict the images in the VDN data set to generate a semantic segmentation mask of the landmark; in the semantic segmentation mask, the pixel value of the landmark area is usually 1 (or a specific value), while the pixel value of other areas is 0; the generated mask is preprocessed such as denoising, morphological operations (dilation, erosion, opening operation, closing operation) to optimize the extraction effect of the bounding box; then the contour detection function in the image processing library (such as OpenCV) is used to detect the contour of the landmark on the preprocessed mask, and the detected contour is fitted with a minimum circumscribed rotated rectangle algorithm (such as OpenCV's minAreaRect function) to fit the rotated bounding box to obtain the center point coordinates, width, height and rotation angle of the bounding box; finally, the language entity and the landmark mask are preliminarily aligned, and the incorrectly matched entity-landmark pairs are corrected through manual verification to obtain refined aerial visual dialogue navigation data.
[0077] The navigation model training method for refined alignment learning provided by an embodiment of the present invention extracts language entities from the navigation dialogue data of the AVDN dataset, obtains landmark masks from the AVDN dataset, and extracts the rotation bounding box information of the landmark masks, performs preliminary alignment on the language entities and landmark masks, and corrects the mismatched entity-landmark pairs through manual verification to obtain refined aerial visual dialogue navigation data. The aerial visual dialogue navigation data containing detailed semantic annotations, landmark geometric features and multimodal information can be obtained, and errors in entity-landmark alignment are corrected through a manual verification mechanism to ensure data diversity and high quality, thereby providing reliable data support for training and evaluating navigation models.
[0078] In some embodiments, after obtaining the aerial visual dialogue navigation model, the method also includes: a navigation strategy model based on time series transformation fuses and makes decisions on the navigation results output by the aerial visual dialogue navigation model to obtain a target navigation strategy; wherein the navigation strategy model based on time series transformation is used to predict the navigation trajectory of the drone based on semantic grid features, existing dialogue data, visual features and trajectory data.
[0079] In this embodiment, in order to cope with multimodal navigation tasks in complex scenarios, a navigation strategy model based on temporal transformation is adopted by combining the dialogue history , Visual History , Track History and semantic grid features , the three-dimensional navigation action of the UAV is generated based on the fusion of multimodal features, so that the trained navigation model can efficiently predict the navigation path in complex scenes, with strong robustness and real-time performance.
[0080] In this embodiment, the navigation strategy model based on temporal transformation is combined with the conversation history , Visual History , Track History Multimodal historical information and semantic grid features , generating navigation actions through a feedforward neural network , the calculation formula is: ; in, Represents the three-dimensional spatial displacement of the UAV; by fusing multimodal features, the model can generate high-precision navigation paths in complex scenes; the navigation model has the ability to efficiently handle multimodal navigation tasks in complex scenes, with both high precision and real-time reasoning performance, and is suitable for a variety of aerial navigation applications.
[0081] The navigation model training method for refined alignment learning provided in an embodiment of the present invention fuses and makes decisions on the navigation results output by the aerial visual dialogue navigation model through a navigation strategy model based on time series transformation, obtains a target navigation strategy, and realizes efficient multimodal information fusion and feature extraction. Compared with the traditional navigation strategy model, the transformer structure better captures long-range dependencies and local context information, thereby generating a more accurate and robust navigation path in complex scenarios.
[0082] Figure 2 This is the second flow chart of the navigation model training method for refined alignment learning provided by the present invention. Figure 2 In the embodiment shown, the ORCNN network is used to obtain aerial images from drones. Extract landmark object encoding (corresponding to semantic features) from The position information of the drone extracted from the image is used to obtain the polar coordinate embedding (corresponding to the spatial features). Perform visual encoding to obtain potential features (corresponding to visual features), then perform weighted calculation on the above three types of features and combine them to obtain semantic grid features , and then use the landmark rotation bounding box prediction task, landmark semantic prediction task and entity-landmark comparison learning task according to Displaying the learned entity-landmark alignment relationship, including using the pre-set FG-AVDN dataset as training samples (such as the input "fly over the C-shaped wheat building and the southernmost white plane to reach the destination"), The jointly constructed comprehensive loss is used to iteratively train the navigation model. The navigation prediction results and the text encoding results corresponding to the input text are input into the navigation strategy model based on time series transformation to complete the fusion and decision-making of multimodal features. Finally, the navigation strategy is generated, that is, the UAV action prediction results and progress prediction results are obtained.
[0083] The navigation method provided by the present invention is described below. The navigation method described below and the navigation model training method for refined alignment learning described above can be referenced to each other.
[0084] Figure 3 It is a flowchart of the navigation method provided by the present invention, such as Figure 3 As shown, the navigation method includes the following steps: Step 310: Obtain the drone aerial image to be processed.
[0085] In this step, the drone aerial images to be processed can be images obtained from a refined aerial visual dialogue navigation dataset or a drone aerial image database, or can be images taken in real time.
[0086] In this embodiment, the drone aerial images to be processed include contents such as airports, drones, buildings, and landmarks.
[0087] Step 320: Process the unmanned aerial vehicle aerial image to be processed based on the aerial visual dialogue navigation model to obtain a navigation result; the aerial visual dialogue navigation model is trained by a navigation model training method of refined alignment learning.
[0088] In this step, the aerial visual dialogue navigation model is trained through the following steps: (1) Visual features, semantic features, and spatial features are extracted from drone aerial images, and weighted fusion is performed on the visual features, semantic features, and spatial features to obtain semantic mesh features. The semantic features are used to represent the category information and geometric information of landmark objects in drone aerial images; the spatial features are used to represent the relative direction and relative distance between the drone and the landmark objects. (2) Based on multiple auxiliary prediction tasks, the refined alignment relationship between entities and landmark objects is explicitly learned according to the semantic grid features to obtain a visual representation; wherein the multiple auxiliary prediction tasks include at least two of the following: a landmark rotation bounding box prediction task, a landmark semantic prediction task, and an entity-landmark comparison learning task; (3) The refined aerial visual dialogue navigation dataset is used as training samples, visual representation is used as input features, and the comprehensive loss is used as the loss function to iteratively train the navigation model to obtain the aerial visual dialogue navigation model; wherein the comprehensive loss is determined based on the navigation loss function and the loss functions corresponding to multiple auxiliary prediction tasks; the refined aerial visual dialogue navigation dataset is obtained by fine-grained expansion of the entity-landmark data in the UAV conversational navigation AVDN dataset.
[0089] It should be noted that the implementation method of each training step corresponds one-to-one to the above-mentioned steps 110 to 130, and will not be repeated in this embodiment.
[0090] Step 330: The navigation strategy model based on time series transformation fuses and makes decisions on the navigation results to obtain a navigation strategy; wherein the navigation strategy model based on time series transformation is used to predict the navigation trajectory of the UAV according to semantic grid features, existing conversation data, visual features and trajectory data.
[0091] In this embodiment, a navigation strategy model based on temporal transformation is used to combine the conversation history. , Visual History , Track History and semantic grid features , based on the fusion of multimodal features, the three-dimensional navigation action of the UAV is generated, so that the trained navigation model can efficiently predict the navigation path in complex scenes; the calculation formula of the predicted path of the UAV is: ; in, Represents the three-dimensional spatial displacement of the UAV; by fusing multimodal features, the model can generate high-precision navigation paths in complex scenes; the navigation model has the ability to efficiently handle multimodal navigation tasks in complex scenes, with both high precision and real-time reasoning performance, and is suitable for a variety of aerial navigation applications.
[0092] In this embodiment, the drone aerial images to be processed include airports, drones, buildings, landmarks and other contents. The drone aerial images to be processed are processed based on the aerial visual dialogue navigation model to obtain the drone navigation strategy, and then the navigation strategy model based on time transformation is used to combine the semantic grid features corresponding to the dialogue history, visual history, trajectory history and drone navigation path A to generate the three-dimensional navigation action of the drone based on multimodal feature fusion, thereby generating a more accurate drone navigation strategy.
[0093] The navigation method provided in the embodiment of the present invention processes the drone aerial images to be processed based on the aerial visual dialogue navigation model to obtain navigation results, and then uses the navigation strategy model based on time series transformation to fuse and make decisions on the navigation results to obtain a navigation strategy, thereby improving the navigation accuracy and robustness of the drone in complex dynamic scenes. The method can be widely used in practical application scenarios such as urban logistics, agricultural monitoring, search and rescue, and industrial inspection.
[0094] The navigation model training device for refined alignment learning provided by the present invention is described below. The navigation model training device for refined alignment learning described below and the navigation model training method for refined alignment learning described above can be referenced to each other.
[0095] Figure 4 : is a structural diagram of a navigation model training device for refined alignment learning provided by the present invention, such as Figure 4 As shown, the navigation model training device for refined alignment learning includes: a feature fusion module 410, an alignment learning module 420 and a training module 430.
[0096] The feature fusion module 410 is used to extract visual features, semantic features and spatial features from the drone aerial images, and perform weighted fusion of the visual features, semantic features and spatial features to obtain semantic mesh features; wherein the semantic features are used to represent the category information and geometric information of the landmark objects in the drone aerial images; and the spatial features are used to represent the relative direction and relative distance between the drone and the landmark objects; An alignment learning module 420 is used to display and learn the refined alignment relationship between the entity and the landmark object according to the semantic grid features based on multiple auxiliary prediction tasks to obtain a visual representation; wherein the multiple auxiliary prediction tasks include at least two of the landmark rotation bounding box prediction task, the landmark semantic prediction task, and the entity-landmark comparison learning task; The training module 430 is used to iteratively train the navigation model using the refined aerial visual dialogue navigation dataset as training samples, visual representation as input features, and comprehensive loss as loss function to obtain an aerial visual dialogue navigation model; wherein the comprehensive loss is determined based on the navigation loss function and the loss functions corresponding to multiple auxiliary prediction tasks; the refined aerial visual dialogue navigation dataset is obtained by fine-grained expansion of the entity-landmark data in the drone conversational navigation AVDN dataset.
[0097] The navigation model training device for refined alignment learning provided by the embodiment of the present invention obtains semantic grid features by weighted fusion of visual features, semantic features and spatial features extracted from aerial images of unmanned aerial vehicles, and uses multiple auxiliary prediction tasks to display and learn the refined alignment relationship between entities and landmark objects according to the semantic grid features to obtain visual representations. Finally, the navigation model is iteratively trained using a refined aerial visual dialogue navigation dataset as training samples, visual representations as input features, and comprehensive loss as the loss function to obtain an aerial visual dialogue navigation model. By comprehensively integrating multimodal features (including visual, semantic and spatial information), the navigation accuracy, alignment capability and task execution efficiency of the unmanned aerial vehicle in complex scenes are improved.
[0098] The navigation device provided by the present invention is described below. The navigation device described below and the navigation method described above can be referred to each other.
[0099] Figure 5 is a schematic diagram of the structure of the navigation device provided by the present invention, such as Figure 5 As shown, the navigation device includes: an image acquisition module 510, a navigation prediction module 520 and a navigation decision module 530.
[0100] An image acquisition module 510 is used to acquire the drone aerial image to be processed; The navigation prediction module 520 is used to process the unmanned aerial vehicle aerial image to be processed based on the aerial visual dialogue navigation model to obtain a navigation result; the aerial visual dialogue navigation model is trained by a navigation model training method of refined alignment learning; The navigation decision module 530 is used to fuse and make decisions on navigation results based on the time-series transformation navigation strategy model to obtain a navigation strategy; wherein the time-series transformation-based navigation strategy model is used to predict the navigation trajectory of the drone based on semantic grid features, existing conversation data, visual features and trajectory data.
[0101] The navigation device provided in the embodiment of the present invention processes the drone aerial images to be processed based on the aerial visual dialogue navigation model to obtain navigation results, and then uses the navigation strategy model based on time series transformation to fuse and make decisions on the navigation results to obtain a navigation strategy, thereby improving the navigation accuracy and robustness of the drone in complex dynamic scenes. The device can be widely used in practical application scenarios such as urban logistics, agricultural monitoring, search and rescue, and industrial inspection.
[0102] Figure 6 is a schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 6As shown, the electronic device may include: a processor (processor) 610 , a communication interface (Communications Interface) 620 , a memory (memory) 630 and a communication bus 640 , wherein the processor 610 , the communication interface 620 , and the memory 630 communicate with each other through the communication bus 640 . The processor 610 can call the logic instructions in the memory 630 to execute the navigation model training method of refined alignment learning, which includes: extracting visual features, semantic features and spatial features from the drone aerial images, and weighted fusion of the visual features, semantic features and spatial features to obtain semantic grid features; wherein the semantic features are used to represent the category information and geometric information of the landmark objects in the drone aerial images; the spatial features are used to represent the relative direction and relative distance between the drone and the landmark objects; based on multiple auxiliary prediction tasks, the refined alignment relationship between the entity and the landmark object is displayed and learned according to the semantic grid features to obtain a visual representation; wherein the multiple auxiliary prediction tasks include at least two of the landmark rotation bounding box prediction task, the landmark semantic prediction task and the entity-landmark comparison learning task; using the refined aerial visual dialogue navigation dataset as a training sample, the visual representation as an input feature, and the comprehensive loss as a loss function to iteratively train the navigation model to obtain an aerial visual dialogue navigation model; wherein the comprehensive loss is determined based on the navigation loss function and the loss function corresponding to the multiple auxiliary prediction tasks; the refined aerial visual dialogue navigation dataset is obtained by fine-grained expansion of the entity-landmark data in the drone conversational navigation AVDN dataset.
[0103] Or execute a navigation method, which includes: obtaining the unmanned aerial vehicle aerial images to be processed; processing the unmanned aerial vehicle aerial images to be processed based on an aerial visual dialogue navigation model to obtain a navigation result; the aerial visual dialogue navigation model is trained by a navigation model training method of refined alignment learning; a navigation strategy model based on time series transformation fuses and makes decisions on the navigation results to obtain a navigation strategy; wherein the navigation strategy model based on time series transformation is used to predict the navigation trajectory of the unmanned aerial vehicle based on semantic grid features, existing dialogue data, visual features and trajectory data.
[0104] In addition, the logic instructions in the above-mentioned memory 630 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0105] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the navigation model training method for refined alignment learning provided by the above methods, the method comprising: extracting visual features, semantic features and spatial features from drone aerial images, and weightedly fusing the visual features, semantic features and spatial features to obtain semantic grid features; wherein the semantic features are used to represent the category information and geometric information of landmark objects in drone aerial images; the spatial features are used to represent the relative direction and relative distance between the drone and the landmark object; based on multiple auxiliary prediction The detection task performs explicit learning of the refined alignment relationship between entities and landmark objects according to the semantic grid features to obtain a visual representation; wherein, the multiple auxiliary prediction tasks include at least two of the landmark rotation bounding box prediction task, the landmark semantic prediction task, and the entity-landmark comparison learning task; the refined aerial visual dialogue navigation dataset is used as training samples, the visual representation is used as input features, and the comprehensive loss is used as the loss function to iteratively train the navigation model to obtain the aerial visual dialogue navigation model; wherein, the comprehensive loss is determined based on the navigation loss function and the loss functions corresponding to the multiple auxiliary prediction tasks; the refined aerial visual dialogue navigation dataset is obtained by fine-grained expansion of the entity-landmark data in the drone conversational navigation AVDN dataset.
[0106] Or execute a navigation method, which includes: obtaining the unmanned aerial vehicle aerial images to be processed; processing the unmanned aerial vehicle aerial images to be processed based on an aerial visual dialogue navigation model to obtain a navigation result; the aerial visual dialogue navigation model is trained by a navigation model training method of refined alignment learning; a navigation strategy model based on time series transformation fuses and makes decisions on the navigation results to obtain a navigation strategy; wherein the navigation strategy model based on time series transformation is used to predict the navigation trajectory of the unmanned aerial vehicle based on semantic grid features, existing dialogue data, visual features and trajectory data.
[0107] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a navigation model training method for refined alignment learning provided by the above methods, the method comprising: extracting visual features, semantic features and spatial features from drone aerial images, and weightedly fusing the visual features, semantic features and spatial features to obtain semantic grid features; wherein the semantic features are used to represent the category information and geometric information of landmark objects in drone aerial images; the spatial features are used to represent the relative direction and relative distance between the drone and the landmark object; based on multiple auxiliary prediction tasks, the semantic grid features are used to identify the entity and the landmark object; The refined alignment relationship of landmark objects is displayed and learned to obtain visual representation; wherein, multiple auxiliary prediction tasks include at least two of the landmark rotation bounding box prediction task, the landmark semantic prediction task and the entity-landmark comparison learning task; the refined aerial visual dialogue navigation dataset is used as training samples, the visual representation is used as input features, and the navigation model is iteratively trained with the comprehensive loss as the loss function to obtain the aerial visual dialogue navigation model; wherein, the comprehensive loss is determined based on the navigation loss function and the loss functions corresponding to multiple auxiliary prediction tasks; the refined aerial visual dialogue navigation dataset is obtained by fine-grained expansion of the entity-landmark data in the drone conversational navigation AVDN dataset.
[0108] Or execute a navigation method, which includes: obtaining the unmanned aerial vehicle aerial images to be processed; processing the unmanned aerial vehicle aerial images to be processed based on an aerial visual dialogue navigation model to obtain a navigation result; the aerial visual dialogue navigation model is trained by a navigation model training method of refined alignment learning; a navigation strategy model based on time series transformation fuses and makes decisions on the navigation results to obtain a navigation strategy; wherein the navigation strategy model based on time series transformation is used to predict the navigation trajectory of the unmanned aerial vehicle based on semantic grid features, existing dialogue data, visual features and trajectory data.
[0109] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0110] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0111] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A navigation model training method for refined alignment learning, characterized in that: include: Extracting visual features, semantic features and spatial features from the drone aerial image, and weighted fusion of the visual features, the semantic features and the spatial features to obtain semantic mesh features; wherein the semantic features are used to represent the category information and geometric information of the landmark object in the drone aerial image; and the spatial features are used to represent the relative direction and relative distance between the drone and the landmark object; Based on multiple auxiliary prediction tasks, the refined alignment relationship between the entity and the landmark object is displayed and learned according to the semantic grid features to obtain a visual representation; wherein the multiple auxiliary prediction tasks include at least two of a landmark rotation bounding box prediction task, a landmark semantic prediction task, and an entity-landmark comparison learning task; The refined aerial visual dialogue navigation dataset is used as a training sample, the visual representation is used as an input feature, and the navigation model is iteratively trained with a comprehensive loss as a loss function to obtain an aerial visual dialogue navigation model; wherein the comprehensive loss is determined based on the navigation loss function and the loss functions corresponding to the multiple auxiliary prediction tasks; the refined aerial visual dialogue navigation dataset is obtained by fine-grained expansion of the entity-landmark data in the UAV conversational navigation AVDN dataset.
2. The navigation model training method for refined alignment learning according to claim 1 is characterized in that: The refined aerial visual dialogue navigation dataset is obtained through the following steps: Extracting language entities from navigation dialogue data of the AVDN dataset; the language entities include a plurality of landmark objects, landmark location descriptions and semantic relationships; Obtain a landmark mask from the AVDN dataset, and extract the rotation bounding box information of the landmark mask, wherein the rotation bounding box information includes the landmark center point, width, height, and rotation angle; The language entity and the landmark mask are preliminarily aligned, and the mismatched entity-landmark pairs are corrected through manual verification to obtain the refined aerial visual dialogue navigation data.
3. The navigation model training method for refined alignment learning according to claim 1 is characterized in that: The multiple auxiliary prediction tasks include the landmark rotation bounding box prediction task, the landmark semantic prediction task and the entity-landmark comparison learning task; The display learning of the refined alignment relationship between the entity and the landmark object based on the semantic grid features based on multiple auxiliary prediction tasks to obtain the visual representation includes: Obtaining landmark rotation bounding box information through the landmark rotation bounding box prediction task according to the multi-head cross attention mechanism and the semantic grid feature; Acquire landmark language description information through the landmark semantic prediction task according to the autoregressive mechanism and the semantic grid features; The entity-landmark contrast learning task is used to contrast and learn the cross-modal features of the entity and the landmark according to the landmark rotation bounding box information and the landmark language description information to obtain the visual representation.
4. The navigation model training method for refined alignment learning according to claim 1 is characterized in that: The comprehensive loss is determined by the following formula: ; in, is the comprehensive loss, is the navigation loss function, is the loss function corresponding to the multiple auxiliary prediction tasks; It is expressed by the following formula: ; in, is the loss function corresponding to the landmark rotation bounding box prediction task, is the loss function corresponding to the landmark semantic prediction task, is the loss function corresponding to the entity-landmark comparison learning task; and is a weight parameter used to balance the loss weight of each task.
5. The navigation model training method for refined alignment learning according to claim 1 is characterized in that: After obtaining the aerial visual dialogue navigation model, the method further includes: The navigation strategy model based on time series transformation fuses and makes decisions on the navigation results output by the aerial visual dialogue navigation model to obtain a target navigation strategy; wherein the navigation strategy model based on time series transformation is used to predict the navigation trajectory of the UAV according to the semantic grid features, existing dialogue data, visual features and trajectory data.
6. A navigation method, characterized in that: include: Obtain the drone aerial images to be processed; The unmanned aerial vehicle aerial image to be processed is processed based on the aerial visual dialogue navigation model to obtain a navigation result; The aerial visual dialogue navigation model is trained by the navigation model training method of refined alignment learning according to any one of claims 1 to 5; The navigation strategy model based on time series transformation fuses and makes decisions on the navigation results to obtain a navigation strategy; wherein the navigation strategy model based on time series transformation is used to predict the navigation trajectory of the drone according to semantic grid features, existing conversation data, visual features and trajectory data.
7. A navigation model training device for refined alignment learning, characterized in that: include: A feature fusion module is used to extract visual features, semantic features and spatial features from the drone aerial images, and perform weighted fusion on the visual features, the semantic features and the spatial features to obtain semantic mesh features; wherein the semantic features are used to represent the category information and geometric information of the landmark objects in the drone aerial images; and the spatial features are used to represent the relative direction and relative distance between the drone and the landmark objects; An alignment learning module, configured to display and learn the refined alignment relationship between entities and landmark objects according to the semantic grid features based on multiple auxiliary prediction tasks to obtain a visual representation; wherein the multiple auxiliary prediction tasks include at least two of a landmark rotation bounding box prediction task, a landmark semantic prediction task, and an entity-landmark comparison learning task; A training module is used to iteratively train the navigation model using a refined aerial visual dialogue navigation dataset as a training sample, the visual representation as an input feature, and a comprehensive loss as a loss function to obtain an aerial visual dialogue navigation model; wherein the comprehensive loss is determined based on the navigation loss function and the loss functions corresponding to the multiple auxiliary prediction tasks; the refined aerial visual dialogue navigation dataset is obtained by fine-grained expansion of the entity-landmark data in the UAV conversational navigation AVDN dataset.
8. A navigation device, characterized in that: include: An image acquisition module is used to acquire the drone aerial images to be processed; A navigation prediction module, used for processing the unmanned aerial vehicle aerial image to be processed based on an aerial visual dialogue navigation model to obtain a navigation result; the aerial visual dialogue navigation model is trained by the navigation model training method of refined alignment learning according to any one of claims 1 to 5; A navigation decision module is used to fuse and make decisions on the navigation results based on a time-series transformation navigation strategy model to obtain a navigation strategy; wherein the time-series transformation-based navigation strategy model is used to predict the navigation trajectory of the drone based on semantic grid features, existing conversation data, visual features and trajectory data.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 6 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Navigation device
CH659320A5
Method for determining pointing deviation of remote sensing satellite on geostationary orbit
CN108364279A
Multi-modal POI feature extraction method and device
CN113032672A
Navigation model training method and device, electronic equipment and storage medium
CN118379563A
Customizable lane biasing for an automated vehicle
US20220371585A1