Unmanned aerial vehicle visual retrieval and instruction feedback method based on multi-modal large model
Through the multimodal large model combined with deep convolutional neural network and Transformer architecture, the efficient integration of drone visual images and natural language is achieved, solving the problem of limited autonomy and adaptability of drones in complex environments, and improving the real-time task execution and multi-task adaptability.
Patent Information
- Application Number
- CN202510458867.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-11
AI Technical Summary
Existing drone visual image processing methods are inefficient in complex environments, difficult to efficiently integrate with natural language generation systems, and cannot generate text instructions that meet actual needs, resulting in limited autonomy and adaptability.
The multimodal large model is used to combine deep convolutional neural networks and Transformer architecture to achieve multimodal feature alignment and context analysis of aerial images and task texts through visual image understanding and natural language generation, and generate real-time and accurate operation instructions.
It significantly improves the autonomy and task execution capabilities of the drone in complex environments, enhances the adaptability and real-timeness of complex scenarios, and supports flexible command generation and dynamic adjustment in multi-task scenarios.
Smart Images

Figure CN120298933A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent perception technology combining drones with large models, and specifically provides a method for drone visual retrieval and instruction feedback based on a multimodal large model. Background Art
[0002] In recent years, drone technology has been widely applied in multiple fields, including logistics distribution, agricultural monitoring, disaster relief, urban planning, and security monitoring. Drones have been widely recognized for their flexible mobility and low-cost advantages when performing tasks in complex environments.
[0003] With the continuous expansion of drone functions, the onboard camera devices can collect a large amount of high-resolution image and video data. However, analyzing and processing these data manually is inefficient, and it cannot meet the actual needs in scenarios with high real-time requirements. Drone vision image processing methods mainly focus on object detection, image classification, and semantic segmentation for specific tasks, but their application scope is limited, and it is difficult to integrate with natural language generation systems efficiently. In addition, when existing systems process dynamic scenes in complex environments, they usually cannot accurately generate text instructions that meet actual needs, resulting in limited adaptability and autonomy of drones when performing complex tasks. To improve the autonomy and intelligence level of drones, there is an urgent need for an efficient and accurate visual image understanding method to extract key information and support the execution of subsequent tasks.
[0004] Currently, large-scale pre-trained vision-language models have made significant progress in the fields of image understanding and natural language processing. These models have powerful feature extraction and semantic understanding capabilities, and can extract complex information from massive data and generate high-quality text descriptions. However, applying these large models to drone scenarios still faces certain difficulties, such as real-time performance, environmental complexity, and limited device resources. Therefore, in drone intelligent perception, how to combine large-scale vision-language models, make full use of rich semantic information in the environment to improve visual image understanding capabilities, and generate text instructions in real-time and efficiently has become a research field worthy of attention and with practical application value. Summary of the Invention
[0005] The present invention aims to combine the latest technologies of visual image understanding and natural language generation to enhance the information perception and task execution capabilities of drones in complex environments. The purpose of the present invention is to load a large-scale pre-trained model, perform multimodal feature alignment on aerial images and task text data, combine context analysis and semantic understanding, realize in-depth parsing and dynamic adaptation of the scene, and provide comprehensive technical support for the autonomous operation of drones in complex environments.
[0006] To achieve the above object, the present invention proposes a method for drone vision retrieval and instruction feedback based on a multi-modal large model, including:
[0007] Step S1 Data acquisition and preprocessing: The camera carried by the drone is used to collect environmental images and video data in real time, and the collected data is denoised, enhanced, and color balanced.
[0008] Step S2 Feature extraction: Use the deep convolutional neural network ResNet (Residual Neural Network) to extract visual features of the target area, including color, shape, texture, etc.
[0009] Step S3 Object detection and segmentation: Use the YOLO (You Only Look Once) model to identify objects in the image, and use the Mask R-CNN (Mask Region-based Convolutional Neural Network) algorithm to segment each region in the image.
[0010] Step S4 Large model loading: Load the pre-trained CLIP (Contrastive Language-Image Pretraining) vision-language model and the T5 (Text-to-Text Transfer Transformer) large language model.
[0011] Step S5 Image and text encoding: Input the preprocessed image into the visual encoder of the CLIP model to generate a high-dimensional feature vector, and encode the predefined and user-provided task descriptions through the text encoder to generate a text feature vector.
[0012] Step S6 Multi-modal alignment: Align the image and text features through the CLIP model and evaluate their correlation.
[0013] Step S7 Context analysis: Adopt the Transformer architecture and combine the time image sequence and task context to improve the model's understanding ability.
[0014] Step S8 Semantic understanding: Adopt the Transformer architecture to perform semantic interpretation on the key content in the visual data.
[0015] Step S9 Instruction generation: The T5 model generates a template according to the text instructions preset by the task objective, and dynamically generates accurate operation instructions based on the image semantic information and context.
[0016] Step S10 Real-time correction and optimization: Adjust and optimize the generated instructions in real time to adapt to environmental changes.
[0017] Step S11 Instruction Output: Output the generated text instructions in an executable format to the UAV control system and return them to the human-machine interaction interface for easy user understanding and adjustment.
[0018] The described camera module is the core component in the UAV system for real-time collection of environmental images and video data. This module includes a high-resolution camera, an infrared sensor, a multispectral camera, an image processing unit, and a communication interface with the UAV main control system. This module is not only responsible for data collection of the UAV but also provides a high-quality data foundation for subsequent analysis with the help of advanced image processing technology.
[0019] The described ResNet deep convolutional neural network solves the problems of gradient vanishing and gradient explosion in network training through residual connections, greatly improving the training stability and convergence speed of deep networks. In the UAV vision system, ResNet can identify high-dimensional features such as the color, shape, and texture of the target area, improving the accuracy of target detection and segmentation.
[0020] The described YOLO model is a real-time object detection algorithm that can quickly and accurately identify objects (such as people, vehicles, buildings, etc.) in images and output the category and bounding box position of the objects, providing real-time environmental perception function for the UAV vision system.
[0021] The described Mask R-CNN model is a deep learning model widely used in segmentation tasks. It can perform semantic segmentation and multi-object instance segmentation on the target area. Semantic segmentation generates a pixel-level segmentation mask for each detected object to obtain the precise contour of the object. Multi-object instance segmentation supports the segmentation of multiple objects in the image and distinguishes different instances of the same category.
[0022] The described CLIP model is trained based on the contrastive learning method and can efficiently perform joint representation learning on images and texts. This model is trained with a large-scale image-text data pair, enabling it to directly understand natural language descriptions and associate corresponding visual information without specific task supervision. In UAV applications, CLIP can be used for image content parsing, converting the visual data captured by the UAV into a high-dimensional feature representation and combining it with natural language descriptions to provide key inputs for subsequent tasks.
[0023] The described T5 model is a large-scale pre-trained generation model integrated in the system. As the core language processing unit, it is responsible for parsing text task objectives, generating precise operation instructions, performing complex task Q&A, and reasoning based on the visual-text descriptions provided by CLIP in a multi-modal environment.
[0024] The CLIP image and text encoding module is responsible for converting the input image and text data into high-dimensional feature representations, enabling the two to be associated and compared in a unified feature space. Image encoding uses a visual encoder to convert the input image into a high-dimensional feature vector, capturing visual features such as color, texture, shape, and spatial information in the image. Text encoding converts the input natural language description into a high-dimensional feature vector through a text encoder, extracting semantic information.
[0025] The CLIP multimodal alignment module aims to map visual features and language features to a unified embedding space, reflecting the semantic consistency between the two in the feature space, and quantifying the degree of association between the two by calculating the similarity of image and text features. By aligning visual features (images) and language features (text), it provides a feature basis for multimodal understanding, generation, and retrieval tasks.
[0026] The described context analysis module is a deep learning component based on the Transformer architecture, aiming to enhance the model's understanding ability of complex scenarios by comprehensively modeling and analyzing task-related context information. This module can effectively capture the logical relationships and cross-modal information in time series data, enabling intelligent systems to have more accurate context awareness capabilities.
[0027] The described semantic understanding module is also a deep learning component based on the Transformer architecture, used to semantically interpret the key content in visual data. This module extracts the deep semantic information of visual data through self-attention mechanisms and feature fusion techniques, and generates semantic expressions that meet the task requirements.
[0028] The described instruction generation module is a core component based on the large-scale pre-trained model T5. Utilizing its powerful natural language processing and generation capabilities, it dynamically generates accurate operation instructions according to task objectives, image semantic information, and context. This module combines preset rules with dynamic understanding of context and can generate task instructions efficiently and flexibly.
[0029] The described real-time correction and optimization module is a dynamic regulation component, aiming to adjust and optimize the generated instructions in real time according to environmental changes and task execution status, ensuring that the system has higher adaptability and response capabilities in complex and dynamic scenarios.
[0030] The described instruction output module is a key component in the system responsible for converting the generated text instructions into an executable format and sending them to the drone control system. This module ensures the accuracy and executability of the instructions, enabling the drone to efficiently complete the specified actions according to the task objectives.
[0031] Compared with the prior art, the beneficial effects of the present invention are:
[0032] First, existing object detection and image understanding technologies are often sensitive to environmental conditions. For example, changes in lighting, interference from complex backgrounds, etc. can easily lead to a decline in model performance. The present invention significantly enhances the adaptability to complex scenarios by combining attention mechanisms, multi-scale feature extraction, and context analysis technologies.
[0033] Second, existing technologies often optimize for specific tasks (such as object detection, path planning), and the applicable scenarios are relatively single. Through the pre-training ability of the CLIP vision-language model and the multi-task learning framework, the present invention supports multiple task scenarios in the same system. This multi-task support ability reduces the cost of developing and deploying different task modules.
[0034] Third, the instructions generated by traditional methods usually mainly rely on fixed templates or limited semantic rules and are difficult to adapt to diverse task requirements. With the natural language generation ability of the T5 large language model, the present invention can generate flexible, diverse, and semantically rich task instructions. This ability greatly improves the practicality of drones in fixed task scenarios.
[0035] Fourth, in existing technologies, the task planning and execution of drones are usually pre-set and lack the ability of real-time dynamic adjustment. By introducing a real-time correction and optimization module and combining the aerial image data of the drone, the present invention can dynamically update instructions according to environmental changes during task execution. This real-time optimization ability significantly improves the success rate of drone tasks, especially in complex tasks such as search and rescue, security monitoring, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The present invention will be further described below in conjunction with the drawings and examples.
[0037] Figure 1 Flow chart of the method for drone visual retrieval and instruction feedback based on a multi-modal large model of the present invention;
[0038] Figure 2 Drone equipped with a camera in the present invention;
[0039] Figure 3 Flow chart of ResNet for extracting features of aerial images in the present invention;
[0040] Figure 4 Flow chart of YOLO for object detection of aerial images in the present invention;
[0041] Figure 5 Flow chart of Mask R-CNN for image segmentation of aerial images in the present invention;
[0042] Figure 6 CLIP model diagram for aligning image and text features in the present invention;
[0043] Figure 7 T5 model diagram for instruction generation in the present invention. Detailed implementation manners
[0044] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0045] As Figure 1 shown, the present invention proposes a method for drone vision retrieval and instruction feedback based on a multimodal large model, including the following steps:
[0046] Step S1 Data acquisition and preprocessing:
[0047] This embodiment relates to the process of a drone collecting and preprocessing image and video data in a complex environment. The drone uses the high-resolution camera, infrared sensor and multispectral camera carried thereon to collect image and video data of a large range of environments in real time, and transmits them to the computing terminal through the wireless communication module, as shown in the physical object. Figure 2 These devices are capable of working under different lighting conditions and can obtain clear image and video data in strong light, weak light and night environments.
[0048] Since the drone may encounter problems such as lighting changes, motion blur, and noise interference during flight, this embodiment uses a variety of image processing technologies to improve the data quality: uses an adaptive filtering algorithm to denoise the collected images and reduce the random noise caused by sensor noise and environmental lighting during flight; applies histogram equalization and Gamma correction image enhancement technologies to optimize the contrast and brightness of the images, making the target areas more clearly distinguishable; adopts the ACE (Automatic Color Equalization) color balance algorithm to adjust the RGB channel values of the images, so that the images under different lighting conditions maintain a consistent color style and improve the accuracy of subsequent feature extraction and target detection. Image acquisition also needs to combine the fusion technology of multiple frames of images to improve the quality of a single frame of image.
[0049] For video stream data, the image resolution is improved through inter-frame alignment and super-resolution reconstruction technologies, and the STDANet (Spatio-Temporal Deformable Attention Network) video deblurring algorithm is used to extract clear pixel information in the video segment, so as to better restore the blurred intermediate frames. Finally, the collected image data needs to be cropped and adjusted, and the resolution is adjusted to 512×512 to adapt to the input format of the neural network for subsequent processing.
[0050] Step S2 Feature extraction:
[0051] This embodiment efficiently extracts features from image data based on a deep learning model to accurately identify the visual information of the target area. The ResNet deep neural network is used to analyze the features of the images collected by the drone, and the process is as follows Figure 3 shown.
[0052] The convolutional layer of ResNet is used to extract low-level features, including edge, corner, and texture information, and deeper networks are constructed through residual blocks to capture high-level semantic features. To improve the robustness of feature extraction, this embodiment introduces a multi-scale feature fusion mechanism, that is, after extracting features in different convolutional layers, scale transformation and fusion are performed to balance the detection requirements of small and large targets.
[0053] To meet the requirements of different task scenarios, the ResNet model is optimized through transfer learning. Using the model weights pre-trained on a large-scale dataset, the specific environment images taken by the drone are fine-tuned to improve the recognition accuracy.
[0054] Step S3 Target detection and segmentation:
[0055] This embodiment combines the YOLO target detection model and the Mask R-CNN image segmentation algorithm to perform target recognition and precise segmentation on the environmental images taken by the drone. The YOLO process is as follows Figure 4 shown, and the Mask R-CNN process is as follows Figure 5 shown.
[0056] The YOLO model globally scans the input image through a single-stage target detection method and predicts the bounding box information of different target classes. Compared with traditional target detection methods (such as R-CNN or Faster R-CNN), YOLO has a significant advantage in detection speed, enabling the drone to quickly identify targets such as people, vehicles, and buildings during real-time flight. At the same time, to improve the detection accuracy, the improved YOLOv11 version is adopted, and the feature pyramid module is optimized in the network structure to enhance the detection ability for small targets.
[0057] After the target detection is completed, this embodiment further uses the Mask R-CNN algorithm for target segmentation to obtain a more accurate target contour. Mask R-CNN adds a fully convolutional network branch on the basis of Faster R-CNN, which can generate pixel-level segmentation masks for each detected target.
[0058] Step S4 Large model loading:
[0059] This embodiment mainly involves the loading and optimization of large-scale pre-trained models to support the visual understanding and text instruction generation of drones in complex scenarios. When a drone executes a task, it needs to process a large amount of high-resolution image and real-time video stream data. Therefore, this method adopts two pre-trained models, CLIP and T5, to generate textual descriptions and instructions through visual data. The CLIP vision-language model is as shown in Figure 6 and the T5 large language model is as shown in Figure 7 .
[0060] Load the CLIP model trained with large-scale image-text data on the cloud server to enable it to have a wide range of visual concept understanding capabilities. The visual encoder of CLIP can parse real-time images from the drone, and the text encoder can process task-related text information, thus forming a unified feature space.
[0061] Load the T5 model in the system as a text generation module to support subsequent natural language generation tasks. Since pre-trained models are usually large, to improve the inference speed and resource utilization rate, this embodiment adopts model INT8 (8-bit Integer) quantization and pruning techniques to reduce the computational complexity.
[0062] During the model loading process, adopt a distributed computing framework to allocate the model inference task to the cloud GPU and the local embedded AI chip, so as to achieve efficient and low-latency model operation.
[0063] Step S5 Image and Text Encoding:
[0064] This embodiment efficiently encodes the image data collected by the drone based on the CLIP model, and combines it with the task description text provided by the user to generate corresponding text feature vectors.
[0065] Input the pre-processed image into the visual encoder of the CLIP model. This encoder is composed of ViT (Vision Transformer) and can map the input image to a high-dimensional feature vector space. This feature vector not only contains the basic visual attributes (color, shape, texture, etc.) of the object, but also can capture the overall semantic information of the scene, such as the building complex, traffic conditions, and natural environment captured by the drone.
[0066] While completing image encoding, the CLIP model also receives task-related text descriptions input by the user, such as "Search for the target vehicle", "Detect the construction area", "Identify abnormal crowd gatherings", etc., and inputs these texts into the text encoder of CLIP. The text encoder adopts the Transformer architecture and can convert the input natural language into corresponding high-dimensional vector representations. To enhance the effect of text encoding, this embodiment supports dynamic task description expansion, that is, based on the user's initial text, the system can automatically supplement additional information in combination with the context to improve the integrity of task expression. For example, in the drone inspection task, when the user inputs "Check the exterior wall of the high-rise building", the system further expands it to "Check whether there are cracks, peeling, and foreign object obstructions on the exterior wall of the high-rise building".
[0067] Step S6 Multimodal Alignment:
[0068] This embodiment uses the CLIP model for multimodal alignment to evaluate the correlation between the images captured by the drone and the task text description, so as to achieve accurate target matching. The core of multimodal alignment is to project the image feature vector and the text feature vector into the same embedding space and calculate the similarity score between the two.
[0069] The CLIP model uses the method of contrastive learning to align the images and correct text descriptions in the same scene, and impose penalties on the mismatched image-text pairs to optimize the representation ability of the model. When the drone transmits new images, the system will automatically retrieve the most matching text description and calculate the matching degree between the image and the task description. For example, in the intelligent inspection application, if the drone captures an image of the bridge structure and the task description is "Check the bridge for cracks", the CLIP model will calculate the similarity between the two and decide whether to trigger further analysis based on the matching degree.
[0070] This embodiment also introduces the STN (Spatial Transformer Networks) attention mechanism, which can enhance the matching accuracy of key regions. For example, if the image contains multiple objects (vehicles, pedestrians, buildings, etc.), the system can calculate through attention weights to make the target region most relevant to the task description obtain a higher matching score, thereby improving the accuracy of recognition.
[0071] Step S7 Context Analysis:
[0072] This embodiment uses the Transformer architecture to deeply analyze the temporal image sequence captured by the drone and the task context to improve the model's understanding ability. In actual application scenarios, drones usually collect data in the form of video streams, and single-frame images may be difficult to provide complete environmental information. Therefore, this method introduces a time series modeling mechanism, uses the Transformer network to process consecutive frame images, and combines task context information to enhance the system's temporal understanding ability.
[0073] The image sequence captured by the drone is input into the temporal Transformer model, which uses the self-attention mechanism to analyze the spatio-temporal correlation between multiple frames of images. For example, in an inspection task, if the drone captures subtle changes in the bridge structure (such as crack expansion or support structure deformation) in multiple consecutive frames, the model can infer the change trend through time series analysis and determine whether the anomaly is gradually intensifying.
[0074] The system also combines task context information, such as the environment (city, forest, sea) where the drone executes the task, weather conditions (sunny, rainy, foggy), and task objectives (disaster monitoring, building inspection, target search, etc.), to provide more targeted analysis capabilities. Through the fusion of context information, the model can adjust the way of interpreting visual data in a specific environment. For example, in a forest fire monitoring task, the system can prioritize key information such as the fire source and the direction of smoke spread, while in a building inspection task, it focuses on analyzing anomalies such as wall cracks and structural damage.
[0075] Step S8 Semantic understanding:
[0076] This embodiment performs deep semantic understanding on the visual data captured by the drone based on the Transformer architecture to extract key content related to the task. In traditional computer vision methods, object detection and segmentation mainly rely on the convolutional neural network CNN, but it is difficult to capture long-range dependencies and global semantic information. Therefore, this method uses the Transformer architecture and uses its global attention mechanism to perform semantic analysis on the important content in the drone image.
[0077] The system receives the preprocessed and multi-modal aligned image features and inputs them into the Transformer network for self-attention calculation to identify the core objects in the scene. For example, when the drone conducts urban road inspections, the Transformer can automatically identify abnormal objects such as road cracks, traffic congestion areas, and illegal buildings, and determine whether these anomalies are new problems based on historical inspection data.
[0078] The Transformer can also automatically screen out the visual information most relevant to the current task objective. For example, in a search and rescue mission, if a drone captures a large area of complex terrain, the system can focus on areas where trapped people may be present, such as near rivers, collapsed areas, or inside high-risk buildings.
[0079] To improve the accuracy of semantic understanding, the model adopts knowledge distillation and self-supervised learning methods, enabling it to autonomously adapt to changes in different task scenarios. For example, in an industrial inspection task, the system can automatically generalize key abnormal patterns by comparing the visual features of normal and abnormal devices and form interpretable inspection results.
[0080] Step S9 Instruction Generation:
[0081] This implementation is based on the T5 large language model. According to the task objective, a text instruction template is preset, and combined with image semantic information and context, dynamic and accurate operation instructions are generated. In traditional drone applications, task instructions are often written manually and generated according to fixed rules, making it difficult to adapt to the dynamic changes of complex environments. This method enables the system to have stronger natural language generation capabilities and output intelligent and personalized instructions by introducing the T5 pre-trained language model.
[0082] The system loads the corresponding instruction template based on the task type and target scenario. For example, in a bridge inspection task, the system can preset instruction templates such as "Detect whether there are cracks or deformations in the bridge structure", while in a disaster monitoring task, templates such as "Identify the flood impact area" or "Determine the fire spread trend" can be adopted.
[0083] The T5 model combines image semantic information and task context to dynamically fill and adjust the template. For example, if the drone detects a crack in the support structure during bridge inspection, the T5 model can automatically generate more specific instructions, such as "There is a crack about 20 cm long in the 3rd support column of the bridge. It is recommended to conduct a detailed inspection".
[0084] The T5 model can also generate a continuous instruction stream based on the context to guide the drone to perform multi-step tasks. For example, in a target search and rescue task, the system can generate subsequent instructions such as "A suspected trapped person is detected in area A. It is recommended to lower the altitude for further confirmation" and "After confirming the target, drop rescue supplies and send location information".
[0085] Step S10 Real-time Correction and Optimization:
[0086] This embodiment performs real-time correction and optimization of the commands generated by the UAV based on the dynamic changes of the environment to improve the accuracy and adaptability of task execution. In actual UAV application scenarios, the environment is often complex and variable. For example, wind speed changes may affect the flight trajectory, light changes may interfere with visual recognition, or the movement of the target object may render the original commands inapplicable. Therefore, the system needs to have an adaptive adjustment ability to ensure that the commands always meet the task requirements. To achieve this goal, this method uses an algorithm that combines a feedback mechanism with reinforcement learning to iteratively optimize the commands by real-time sensing of environmental information.
[0087] The system continuously collects sensor data during the UAV's task execution, including GPS position information, IMU inertial data, visual feedback, task completion status, etc., and inputs this data into the optimization module for analysis. For example, in an inspection task, if the UAV flies to the target area according to the preset commands but is affected by strong winds and causes the flight path to deviate, the system will automatically adjust the flight path and correct the subsequent commands to adapt to the new position information.
[0088] In a target detection task, if the initially generated commands require the UAV to focus on a specific area, but as the UAV approaches the target, the system finds that there is actually no abnormality in this area, it will adjust the commands in real-time and shift the focus to a more potentially risky area, such as a neighboring damaged structure or equipment. In addition, to improve the optimization efficiency, this method combines reinforcement learning and continuously improves the command generation strategy through policy gradient optimization, making the commands more robust in complex environments.
[0089] Step S11 Command output:
[0090] This embodiment involves converting the optimized text commands into an executable format and outputting them to the UAV control system, while also feeding them back to the human-machine interface to achieve intelligent command management and adjustment. When the UAV executes a task, it usually needs to operate based on high-level text commands, while low-level control commands (adjusting flight speed, changing course, starting specific sensors, etc.) are executed by the flight control system. Therefore, this method needs to implement the conversion from high-level text commands to executable control commands.
[0091] The system inputs the optimized text commands into the command parsing module, which uses a semantic parsing algorithm based on natural language processing to structurally disassemble the commands and match the corresponding flight control command set. For example, if the text command is "hover over the target area and take high-definition images", the parsing module decomposes it into a series of control commands, including adjusting the UAV's altitude, setting the hovering radius, starting the camera, and adjusting the shooting angle. At the same time, for different task requirements, the system generates command outputs in multiple formats, with the JSON format for API interface calls and binary commands for directly controlling the flight control system.
[0092] This embodiment also returns the instructions to the human-machine interaction interface for easy understanding and adjustment by the operator. During the inspection task, the user can view the instructions generated by the system, such as "Cracks are detected on the surface of the transmission tower. Please confirm further", and manually correct the detection area or adjust the task priority. In addition, to improve the user experience, the system can provide a visual interaction method, mark the drone flight path on the map interface, and allow the user to adjust the task position and parameters by dragging.
Claims
1. A method for visual retrieval and instruction feedback of unmanned aerial vehicles based on multi-modal large models, characterized in that The steps include: collecting environmental images and video data in real time through a camera carried by a drone, and performing noise reduction, enhancement, and color balance processing on the collected image data to improve data quality; using the deep neural network ResNet (Residual Neural Network) to analyze the target regions in the images and extract their visual features, including color, shape, and texture, etc.; using the YOLO (You Only Look Once) model to perform real-time detection of targets (such as people, vehicles, buildings, etc.) in the images, and precisely segmenting each region in the images through the MaskR-CNN (Mask Region-based Convolutional Neural Network) algorithm to extract the boundary and shape information of the target regions; loading large-scale pre-trained models CLIP (Contrastive Language-Image Pretraining) and T5 (Text-to-Text Transfer Transformer); inputting the preprocessed images into the visual encoder of the CLIP model to generate high-dimensional image feature vectors, and at the same time, generating text feature vectors by inputting predefined and user-provided task description texts; achieving multi-modal alignment of image and text features through the CLIP model, evaluating their correlation, and providing relevant information for image understanding and task execution; combining the temporal image sequence and task context, and using a model based on the Transformer architecture to analyze multiple frames of images to enhance the model's ability to understand dynamic changes in the scene; performing semantic interpretation on the key content in the visual data based on the Transformer architecture to extract high-level semantic information; generating an initial text instruction template according to the task objective through the T5 model, and dynamically generating precise operation instructions based on the image semantic information and context; During the task execution process, adjust and optimize the generated instructions in real time to ensure that the operation instructions can adapt to environmental changes and maintain accuracy; output the optimized text instructions in an executable format to the drone control system and synchronously return them to the human-computer interaction interface for the user to view, understand, and adjust.
2. The method for drone vision retrieval and instruction feedback based on a multimodal large model according to claim 1, characterized in that: Collect environmental images and video data in real time through a high-resolution camera carried by a drone, covering various lighting, weather, and complex scene conditions. After collection, perform a series of preprocessing operations on the image data, including denoising to eliminate sensor noise and background interference, image enhancement to improve detail clarity, and color balance to correct image color bias, ensuring that the input data can serve as a reliable basis for subsequent feature extraction and analysis.
3. The method for visual retrieval and instruction feedback of an unmanned aerial vehicle based on a multimodal large model according to claim 1, wherein: The deep neural network ResNet is used to extract features from the images captured by the drone, extracting multi-level visual features including color, shape, texture, etc. The ResNet model can effectively alleviate the problem of gradient disappearance in the training of deep networks through residual connections, improving the feature extraction ability, so as to accurately identify key targets in the image under complex environments. Compared with the traditional CNN architecture, ResNet can capture higher-order feature information of objects more effectively, making the recognition of targets more accurate.
4. The method for visual retrieval and instruction feedback of an unmanned aerial vehicle based on a multimodal large model according to claim 1, characterized in that: Object detection and region segmentation are achieved through deep learning models. In the detection stage, the YOLO model is used to locate key targets in the image in real time, such as pedestrians, vehicles, buildings, etc., and efficient detection is achieved through its fast and high-precision characteristics. In the segmentation stage, the Mask R-CNN model is used to finely segment the target area, extracting boundary, shape and pixel-level area information, providing structured data support for subsequent multi-modal alignment and semantic understanding.
5. The method for visual retrieval and instruction feedback of an unmanned aerial vehicle based on a multimodal large model according to claim 1, wherein: By loading the pre-trained large-scale vision-language model CLIP, the processed image is input into the vision encoder to generate high-dimensional feature vectors. At the same time, according to the task requirements, predefined text and the descriptions provided by the user are input to generate corresponding text feature vectors. Through the multi-modal alignment mechanism of CLIP, the similarity and correlation degree between the image and text features are evaluated to ensure the accuracy of feature alignment. This method can not only handle simple image-text matching tasks, but also support the deep fusion of complex semantic information.
6. The method for drone vision retrieval and instruction feedback based on a multimodal large model according to claim 1, wherein: A deep learning architecture based on Transformer is adopted to combine image sequence data in the time dimension, dynamically capture the temporal change information of the scene, and integrate the task context to improve the model's understanding ability of complex environments. The module can perform semantic interpretation on key targets in the visual data, such as the functions, behavioral characteristics of the targets and their interaction relationships with the environment. At the same time, the module uses context information to analyze the dynamic changes of the environment in real time, providing accurate semantic support for the instruction generation module to ensure that the output operation instructions are closely related to the current environment.
7. The method for drone vision retrieval and instruction feedback based on a multimodal large model according to claim 1, wherein: Relying on the T5 model, a text instruction template based on the task target is generated, and precise operation instructions are dynamically generated in combination with the semantic analysis results. The generation process not only considers the attributes of the target object, but also fully adapts to the dynamic changes of the environment. The real-time correction mechanism optimizes the generated instruction content by analyzing the latest image and context, avoiding incorrect execution or deviation from the task target. The finally output instructions are transmitted to the drone control system in an executable format and displayed on the interaction interface for the user to refer to and adjust, ensuring the transparency, reliability and flexibility of the instructions.
Citation Information
Cited By
Monitoring and end-side analysis method based on long-endurance sounding system
CN120764861A
Multi-level industrial unmanned aerial vehicle control system and method
CN120972747A
Identification information content intelligent detection method based on visual ergonomics
CN121281035A
Visual language model-based intelligent evaluation system and method for rice and fish planting state
CN121544955A
Intelligent evaluation system and method for rice-fish breeding state based on visual language model
CN121544955B