Water conservancy unmanned aerial vehicle inspection method based on visual language action multi-modal model
By processing images and task instructions through a visual language-action multimodal model, the problem of insufficient real-time response capability of the drone inspection system to environmental changes is solved, and efficient execution and intelligent decision-making of water conservancy inspection tasks are achieved.
Patent Information
- Application Number
- CN202510692153.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-26
AI Technical Summary
Existing drone inspection systems lack the ability to respond to environmental changes in real time, have low levels of intelligent decision-making, and are unable to fully utilize multimodal information for adaptive adjustments, resulting in insufficient inspection efficiency and accuracy.
A water conservancy UAV inspection method based on a visual language action multimodal model is adopted. By combining a visual encoder, a language model, a state encoder and an action sequence decoder, image data and task instructions are processed in real time, and action instructions for the UAV and gimbal are generated to achieve understanding and decision-making of complex tasks.
It improves the drone's ability to understand and make real-time decisions in complex tasks, and enables automatic planning of flight paths and efficient execution of water conservancy inspection tasks.
Smart Images

Figure CN120704350A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of drone control technology, and in particular to a water conservancy drone inspection method based on a visual language action multimodal model. Background Art
[0002] With the development of intelligent technologies, particularly in artificial intelligence and automation, drones are increasingly being used, particularly in areas like water conservancy inspections. Traditional inspection methods are not only inefficient but also pose safety risks and are difficult to adapt to complex and dynamic environments. Drones, with their efficient data collection and real-time monitoring capabilities, can significantly improve the efficiency and safety of inspections.
[0003] Current drone inspection systems typically rely on preset flight routes and fixed mission processes, lacking the ability to respond to environmental changes in real time. This design often leads to reduced inspection efficiency and insufficient data collection when faced with complex water conservancy facilities and sudden environmental changes. In addition, the existing system's intelligence level in the decision-making process is low, and it is unable to fully utilize real-time data for adaptive adjustments. Although methods have begun to use small visual perception neural network models such as target detection and semantic segmentation to improve the inspection capabilities of drones, there are still significant deficiencies. For example, the lack of intelligent decision-making capabilities means that existing systems are often unable to make effective decisions based on real-time environmental changes when handling complex tasks; the lack of information fusion means that existing technologies lack the fusion of multimodal information such as the drone's vision, body state, and language tasks, resulting in insufficient accuracy and adaptability of the model when understanding inspection tasks. Summary of the Invention
[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present invention provides a water conservancy drone inspection method based on a visual language action multimodal model.
[0005] In a first aspect, the present invention provides a water conservancy drone inspection method based on a visual language action multimodal model, comprising: S2. In the autonomous inspection mode, the images captured by the drone gimbal camera, the status of the drone flight control and gimbal, and the task instructions are input into the visual language action multimodal model deployed on the drone's onboard computing platform for initialization, and the action that the flight control and gimbal should perform at the next moment is obtained; S3. The drone executes the action instructions of the flight control and gimbal to reach a new state, keeping the two inputs of image and task instructions unchanged, only updating the drone state, and then the visual language action multimodal model predicts the action that the flight control and gimbal should perform at the next moment; S4. Repeat step S3 N times; S5. The image captured by the drone gimbal at the latest moment, the task instructions for the next moment predicted by the visual language action multimodal model, and the latest state of the drone are input into the visual language action multimodal model, and the action of the flight control and gimbal at the next moment is output; Steps S3 to S5 are repeated until the inspection task is completed, and then the drone returns to land; The vision-language-action multimodal model includes: a visual encoder, a language model, a state encoder, an action sequence decoder and an instruction decoder; the visual encoder supports extracting image features from images captured by the drone gimbal camera; the language model supports encoding task instructions and extracting word embeddings of task instructions as instruction text features; the state encoder supports extracting state features from the drone and gimbal state vectors; the action sequence decoder generates the action sequence of the drone and gimbal based on image features, instruction text features and state features, and the instruction decoder generates the instructions that the drone should execute at the next moment based on image features, instruction text features and state features.
[0006] Furthermore, the working process of the visual encoder includes: Use the PatchEmbedding method to preprocess the images collected by the UAV gimbal to obtain feature embedding; The visual encoder uses the ViT model to analyze feature embedding and finally output visual features ; Among them, the ViT model is composed of a combination of multi-layer stacked attention and feedforward neural networks; each layer structure is a series of multi-head attention layer, layer normalization operation, feedforward fully connected layer, and layer normalization operation.
[0007] Furthermore, the PatchEmbedding method is used to preprocess the images collected by the UAV gimbal to obtain feature embedding, including: The visual encoder receives images collected by the drone’s gimbal; Divide the image into multiple fixed-size regions. Assume that the region size of each region is (patch_size, patch_size). Extract all regions from the image through a convolution operation with the convolution kernel size as the region size and the step size as patch_size. After the region is divided, a convolutional layer is used to extract the features of each region at the same time. The output of the convolutional layer generates a new tensor representing the regional features. The tensor extracted by the convolutional layer is flattened to merge the spatial dimensions; then, the flattened tensor is transposed; The position encoding is added to the region features to obtain feature embedding.
[0008] Furthermore, the position encoding adopts relative position encoding: assuming that an input image is divided into regions, for each pair of regions , its relative position encoding is expressed as: ; in, Is used to generate and distance Related encoding functions, This is done using the sine and cosine functions: ; in, is a hyperparameter that controls the frequency of the function. , ; Is a function based on the region position, used to represent the patch and patch The relative position features of : ; in is a learnable weight matrix; All relative position encodings are accumulated to obtain the final output of the position encoding: .
[0009] Furthermore, the language model adopts the 2B Gemma language model in the PaliGemma VLM model.
[0010] Furthermore, the state encoder maps the input drone and gimbal state vectors to a high-dimensional space through a fully connected layer to increase the expressive power of the features, and then uses The activation function introduces nonlinear transformation, and the final output state characteristics are as follows: ; in, is the input drone and gimbal state vector, 、 is the weight matrix and bias vector of the first layer of the fully connected layer, 、 is the weight matrix and bias vector of the second layer of the fully connected layer, It is the final state characteristic; The drone state vector is represented as: ; in, is the drone state vector, 、 、 is the location of the drone, 、 、 is the component speed of the drone, is the heading angle of the drone, is the drone battery level; The gimbal state vector is expressed as: ; in, is the gimbal state vector, is the pitch angle of the gimbal camera, is the roll angle of the gimbal camera point, is the yaw angle of the gimbal camera, is the zoom factor of the gimbal camera.
[0011] Furthermore, the instruction decoder first processes the visual features , the command text features of task instructions and state characteristics The concatenation is performed, and then the concatenated features are converted into the probability distribution of instructions through the linear layer and Softmax, thereby outputting the instructions that should be executed at the next moment. The calculation process is as follows: ; in, Indicates that the visual features , instruction text features and state characteristics Splicing, is the weight matrix of the linear layer of the instruction decoder, is the bias matrix of the linear layer of the instruction decoder.
[0012] Furthermore, the action sequence decoder is based on the Transformer architecture. The input of the action sequence decoder includes: visual features , instruction text features , state characteristics The action sequence decoder consists of a linear layer, a stack of three multi-head attention layers, a linear layer connected to the attention layer to predict the next action of the drone, and a linear layer connected to the attention layer to predict the next action of the gimbal. When predicting actions, the visual features are maintained. Instruction text features No change, only update the drone status characteristics according to the latest drone status , iterate N times and output the drone actions and gimbal actions at N future moments.
[0013] Furthermore, basic natural language task instructions are designed based on the basic capabilities required for drone inspections in water conservancy scenarios. The task instructions include: "continuous flight inspection", "hovering", "flying around the target", "taking accurate photos of the target", "following the target", "flying along the target object", and "pan-tilt tracking the target".
[0014] In the second aspect, the present invention provides a water conservancy drone inspection device based on a visual language action multimodal model, comprising: at least one processing unit, the processing unit and the storage unit are connected through a bus unit, the storage unit is a computer-readable storage medium, and can be used to store software programs, computer executable programs and modules. The processing unit implements the water conservancy drone inspection method based on the visual language action multimodal model by running the software programs, computer executable programs and modules stored in the storage unit.
[0015] In a third aspect, the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed, it implements the water conservancy drone inspection method based on the visual language action multimodal model.
[0016] The above technical solution provided by the embodiment of the present invention has the following advantages compared with the prior art: In scenarios such as drone onboard computing platforms where computing resources are limited and power consumption is extremely sensitive, the present invention provides a lightweight method for action generation. It can reuse image features and instruction text features to generate low-level control actions and achieve real-time control of flight control and gimbal.
[0017] The visual language action multimodal model proposed in this invention can simultaneously process image data and task instructions, achieve information complementarity, and enhance the drone's ability to understand complex tasks. On this basis, it makes real-time decisions and outputs low-level drone control instructions based on multimodal information, realizing flight control and gimbal control of the drone, enabling the drone to automatically plan flight paths and perform water conservancy inspection tasks.
[0018] This paper proposes a visual-language-action multimodal model for water conservancy inspections. This model can be deployed on an unmanned aerial vehicle (UAV) onboard computing platform to generate real-time UAV and gimbal motion and task instructions based on the UAV's visual imagery, task instructions, and UAV status. The design of this model proposes a method for fusing the different modalities of information—the UAV's visual imagery, natural language task instructions, and UAV status. Three encoders process the data for each modality separately and then align them. This achieves a comprehensive understanding of water conservancy tasks, helping to guide the generation of action and task instructions. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0020] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0021] Figure 1 A flowchart of a water conservancy drone inspection method based on a visual language action multimodal model provided by an embodiment of the present invention; Figure 2 This is an architectural diagram of the visual language-action multimodal model provided by an embodiment of the present invention; Figure 3 A flowchart of using the PatchEmbedding method provided in an embodiment of the present invention to preprocess images collected by a drone gimbal to obtain feature embedding; Figure 4 Schematic diagram of a water conservancy drone inspection device based on a visual language action multimodal model provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0023] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0024] Example 1 like Figure 1 As shown, the water conservancy drone inspection method based on the visual language action multimodal model provided by the present invention applies the visual language action multimodal model to the inspection task of the drone water conservancy scene, including: S1. For a drone inspection mission in a water conservancy scenario, the pilot or automated program first controls the drone to fly to a designated location and begin autonomous inspection.
[0025] S2. In autonomous inspection mode, the images captured by the drone's gimbal camera, the status of the drone's flight control and gimbal, and the mission instruction "continuous flight inspection" are input into the visual language action multimodal model deployed on the drone's onboard computing platform for initialization, and the flight control and gimbal actions at the next moment are output.
[0026] S3. The drone executes the flight control and gimbal motion instructions to reach a new state. Only the drone state code is updated. Then, the visual language action multimodal model predicts the action that the flight control and gimbal should perform at the next moment.
[0027] S4. Repeat step S3 N times. N should not be too large to cause information delay. Set N=5.
[0028] S5. Input the latest image captured by the drone gimbal, the next moment’s mission instructions predicted by the visual language action multimodal model, and the latest state of the drone into the visual language action multimodal model, and output the next moment’s flight control and gimbal actions.
[0029] Repeat steps S3 to S5 until the inspection mission is completed, and then return to land.
[0030] This technical solution involves a multi-rotor drone and an onboard computing platform capable of deploying a multimodal vision-language-action model, with a communication link between the two. The multi-rotor drone executes flight commands and gimbal control instructions, while the onboard computing platform, running the multimodal vision-language-action model, calculates the sequence of action commands the drone should execute in the future based on the drone's real-time input of images, language commands, and state.
[0031] The overall architecture of the UAV visual language action multimodal model for water conservancy scene inspection is shown in the attached figure. Figure 2 As shown. The visual language action multimodal model includes: a visual encoder, a language model, a state encoder, an action sequence decoder and an instruction decoder; the visual encoder supports extracting image features from images captured by the drone gimbal camera; the language model supports encoding task instructions and extracting word element embeddings of task instructions as instruction text features; the state encoder supports extracting state features from the drone and gimbal state vectors; the action sequence decoder generates the action sequence of the drone and gimbal based on image features, instruction text features, and state features, and the instruction decoder generates the instructions that the drone should execute at the next moment based on image features, instruction text features, and state features. In the specific implementation process, the next action and task execution are output in the form of returning the probability distribution of the predicted action and task instructions.
[0032] The principles of each part of the visual language action multimodal model are as follows: The visual encoder is used to extract features of images collected by the drone. The working process of the visual encoder includes: Use the PatchEmbedding method to preprocess the images collected by the UAV gimbal to obtain feature embedding, such as Figure 3 As shown, the specific process is as follows: Input image: The visual encoder receives images captured by the drone gimbal.
[0033] Region partitioning: Divide the image into multiple fixed-size regions. Assuming the size of each region is (patch_size, patch_size), all regions are extracted from the image through a convolution operation with the kernel size equal to the region size and a step size of patch_size. The region partitioning process can be thought of as converting the spatial information of the image into a series of small patches, which are further processed in subsequent steps.
[0034] Image region feature extraction: After region division, a convolutional layer is used to simultaneously extract features of each region. The output of the convolutional layer generates a new tensor representing the region features.
[0035] Flattening and transposing: The tensor extracted by the convolutional layer is flattened to merge the spatial dimensions. Then, the flattened tensor is transposed so that subsequent processing steps can easily access the regional features in the tensor.
[0036] Adding position coding: In order to preserve the position information of the region in the original image, position coding is added. When adding, the position coding is added to the regional features so that the model can understand the position of each regional feature and thus capture the spatial structure of the image. The present invention provides a relative position coding method that not only preserves the position information between regional features, but also enhances the model's ability to understand the local context in the image. Assume that there is an input image divided into For each pair of regions , its relative position encoding is expressed as: ; in, Is a function that generates the distance Related coding. Use sine and cosine functions to implement: ; It is a hyperparameter that controls the frequency of the function and helps adjust the oscillation frequency of the sine and cosine functions, thereby affecting the details and range of position encoding. ,in, is the maximum distance of the region: .
[0037] Is a function based on the region position, used to represent the patch and patch For example, using linear transformation modeling : ; in Is a learnable weight matrix with shape (2, d).
[0038] All relative position encodings are accumulated to obtain the final output of the position encoding: ; Finally, PatchEmbedding outputs a feature embedding that fuses regional features and position encoding, which will be used as input for subsequent processing.
[0039] The visual encoder uses the ViT model to analyze feature embedding and finally output visual features The ViT model is composed of a combination of multi-layer stacked attention and feedforward neural networks; each layer is composed of a multi-head attention layer, a layer normalization operation, a feedforward fully connected layer, and a layer normalization operation in series.
[0040] The language model encodes the mission instructions of the UAV and extracts the features of the instruction text. Because training a language model, especially a large multimodal language model that supports processing vision and text, requires a lot of computing power and data, the present invention uses a pre-trained language model as part of the overall architecture to avoid building from scratch. Taking into account the computing power of the UAV, the real-time requirements for predicting mission instructions, and the purpose of the model, the language model of the present invention adopts the 2B Gemma language model in the PaliGemmaVLM model. First, the mission instructions sent to the multi-rotor UAV are segmented, and then the mission instructions are encoded through the 2B Gemma model to obtain the tokens of the mission instructions. Tokens refer to symbols used in the language model to represent Chinese characters, English words, or Chinese and English phrases, and the embedding of the tokens is obtained. The word embedding is a fixed-dimensional vector representation obtained by word mapping, which contains semantic and contextual information. The 2B Gemma language model in the PaliGemma VLM model is a multimodal model that supports processing image and language tasks, receiving feature representations after visual and language alignment, and adopts a decoder-only Transformer architecture. The 2B Gemma language model in the PaliGemma VLM model has low computational complexity and parameter count, runs fast, can adapt to the computing power of drones, and meet the real-time control requirements.
[0041] The state encoder encodes the drone state vector and the gimbal state vector to obtain state features. The state encoder inputs the drone state vector and the gimbal state vector, and converts the state vector into a state code with a unified feature dimension through a fully connected layer.
[0042] The drone state vector is represented as: ; in, is the drone state vector, 、 、 is the location of the drone, 、 、 is the component speed of the drone, is the heading angle of the drone, It is the battery level of the drone.
[0043] The gimbal state vector is expressed as: ; in, is the gimbal state vector, is the pitch angle of the gimbal camera, is the roll angle of the gimbal camera point, is the yaw angle of the gimbal camera, is the zoom factor of the gimbal camera.
[0044] The state encoder encodes the state vectors of the drone and gimbal into state features of a unified feature dimension for subsequent decision making and action generation. The state encoder first performs a linear transformation and maps the input drone and gimbal state vectors to a high-dimensional space through a fully connected layer to increase the feature expression capability, and then uses The activation function introduces nonlinearity, and the final output state characteristics are as follows: ; in, is the input drone and gimbal state vector, 、 is the weight matrix and bias vector of the first layer of the fully connected layer, 、 is the weight matrix and bias vector of the second layer of the fully connected layer, It is the final state characteristic.
[0045] Drone action generation and command decoder generates action sequences and the commands that the drone should execute at the next moment based on multimodal feature representation.
[0046] In this application, task instructions in the form of basic natural language are designed based on the basic capabilities required for drone inspections in water conservancy scenarios. The task instructions include: "continuous flight inspection", "hover", "fly around the target", "take accurate photos of the target", "fly along the target", "fly along the target", and "pan-tilt tracking target". The explanations of the task instructions are as follows:
[0047] The input to the instruction decoder is the visual features , the command text features output after the task command is input into the language model and state characteristics The instruction decoder first concatenates the three input features mentioned above, and then converts the concatenated features into the probability distribution of task instructions through a linear layer and Softmax, thereby outputting the task instruction to be executed at the next moment. The calculation process is as follows: .
[0048] in, Indicates that the visual features , instruction text features and state characteristics Splicing, is the weight matrix of the linear layer of the instruction decoder, is the bias matrix of the linear layer of the instruction decoder.
[0049] Instruction decoder based on visual features , the command text features output after the task command is input into the language model and state characteristics Given the probability distribution of all task instructions, the task instruction with the largest probability distribution is selected as the task instruction to be executed at the next moment.
[0050] The action sequence decoder is based on the Transformer architecture. The input of the action sequence decoder is consistent with the input of the instruction decoder, that is, the visual features , instruction text features , state characteristics The action sequence decoder consists of a linear layer, three stacked multi-head attention layers, a linear layer connected to the attention layer to predict the next action of the drone, and a linear layer connected to the attention layer to predict the next action of the gimbal. , instruction text features No change, only update the drone status characteristics according to the latest drone status , iterate N times and output the drone actions and gimbal actions at N future moments.
[0051] This paper proposes a visual-language-action multimodal model for water conservancy inspections. This model can be deployed on an unmanned aerial vehicle (UAV) onboard computing platform to generate UAV and gimbal motion and task instructions in real time based on the UAV's visual imagery, task instructions, and UAV status. The design of this model incorporates a method for fusing the different modalities of information—the UAV's visual imagery, natural language task instructions, and UAV status. Three encoders process the data from each modality separately and then align them. This enables a comprehensive understanding of water conservancy tasks.
[0052] In scenarios such as drone onboard computing platforms with limited computing resources and extremely sensitive to power consumption, a lightweight method for generating action commands is proposed. It can reuse visual and language features and real-time drone status information to generate low-level control actions and achieve real-time control of flight control and gimbal.
[0053] Example 2 See Figure 4As shown, an embodiment of the present invention provides a transmission channel scene image depth estimation device based on transmission channel geometric prior, including: at least one processing unit, which connects the processing unit and the storage unit through a bus unit. The storage unit is a computer-readable storage medium that can be used to store software programs, computer executable programs and modules, such as the software programs, computer executable programs and modules corresponding to the water conservancy drone inspection method based on the visual language action multimodal model in the embodiment of the present invention. The processing unit implements the above-mentioned water conservancy drone inspection method based on the visual language action multimodal model by running the software programs, computer executable programs and modules stored in the storage unit, including: S2. In the autonomous inspection mode, the images captured by the drone's gimbal camera, the status of the drone's flight control and gimbal, and the task instructions are input into the visual language action multimodal model deployed on the drone's onboard computing platform for initialization, and the flight control and gimbal actions at the next moment are obtained; S3. The drone executes the flight control and gimbal action instructions to reach a new state, only updating the drone's state, and then the visual language action multimodal model predicts the flight control and gimbal actions at the next moment; S4. Repeat step S3N times; S5. The latest image captured by the drone's gimbal, the task instructions for the next moment predicted by the visual language action multimodal model, and the latest state of the drone are input into the visual language action multimodal model, and the flight control and gimbal actions at the next moment are output; Steps S3 to S5 are repeated until the inspection task is completed, and then the drone returns to land; Among them, the visual language action multimodal model includes: a visual encoder, a language model, a state encoder, an action sequence decoder and an instruction decoder; the visual encoder supports extracting image features from images collected by the drone gimbal camera; the language model supports encoding task instructions and extracting word element embeddings of task instructions as instruction text features; the state encoder supports extracting state features from the drone and gimbal state vectors; the action sequence decoder generates the action sequence of the drone and gimbal based on image features, instruction text features, and state features, and the instruction decoder generates the instructions that the drone should execute at the next moment based on image features, instruction text features, and state features.
[0054] Of course, the computer program stored in the storage unit of the transmission channel scene image depth estimation device based on the transmission channel geometric prior provided in an embodiment of the present invention is not limited to the method operations described above, and can also execute related operations in the water conservancy drone inspection method based on the visual language action multimodal model provided in any embodiment of the present invention.
[0055] Example 3 An embodiment of the present invention provides a computer-readable storage medium storing a computer program. When the computer program is executed, the method for water conservancy drone inspection based on a visual language action multimodal model is implemented, including: S2. In the autonomous inspection mode, the images captured by the drone gimbal camera, the status of the drone flight control and gimbal, and the task instructions are input into the visual language action multimodal model deployed on the drone's onboard computing platform for initialization, and the action that the flight control and gimbal should perform at the next moment is obtained; S3. The drone executes the action instructions of the flight control and gimbal to reach a new state, keeping the two inputs of image and task instructions unchanged, only updating the drone state, and then the visual language action multimodal model predicts the action that the flight control and gimbal should perform at the next moment; S4. Repeat step S3 N times; S5. The image captured by the drone gimbal at the latest moment, the task instructions for the next moment predicted by the visual language action multimodal model, and the latest state of the drone are input into the visual language action multimodal model, and the action of the flight control and gimbal at the next moment is output; Steps S3 to S5 are repeated until the inspection task is completed, and then the drone returns to land; Among them, the visual language action multimodal model includes: a visual encoder, a language model, a state encoder, an action sequence decoder and an instruction decoder; the visual encoder supports extracting image features from images collected by the drone gimbal camera; the language model supports encoding task instructions and extracting word element embeddings of task instructions as instruction text features; the state encoder supports extracting state features from the drone and gimbal state vectors; the action sequence decoder generates the action sequence of the drone and gimbal based on image features, instruction text features, and state features, and the instruction decoder generates the instructions that the drone should execute at the next moment based on image features, instruction text features, and state features.
[0056] A computer-readable storage medium provided by an embodiment of the present invention stores a computer program that is not limited to the method operations described above, but can also execute related operations in a water conservancy drone inspection method based on a visual language action multimodal model provided by any embodiment of the present invention.
[0057] In the embodiments provided by the present invention, it should be understood that the disclosed structures and methods can be implemented in other ways. For example, the structural embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, structure or unit, which can be electrical, mechanical or other forms.
[0058] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0059] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0060] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is intended to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A water conservancy drone inspection method based on a visual language action multimodal model, characterized by: include: S2. In autonomous inspection mode, the images captured by the drone's gimbal camera, the status of the drone's flight control and gimbal, and the mission instructions are input into the visual language action multimodal model deployed on the drone's onboard computing platform for initialization. This model then determines the action that the flight control and gimbal should perform at the next moment. S3. The drone executes the flight control and gimbal action instructions to reach a new state, maintaining the image and mission instruction inputs unchanged. Only the drone's state is updated. The visual language action multimodal model then predicts the action that the flight control and gimbal should perform at the next moment. S4. Repeat step S3 N times. S5. Input the latest image captured by the drone gimbal, the next moment's mission instructions predicted by the visual language action multimodal model, and the latest state of the drone into the visual language action multimodal model, and output the next moment's flight control and gimbal actions. Repeat steps S3-S5 until the inspection mission is completed, and then return to land. The vision-language-action multimodal model includes: a visual encoder, a language model, a state encoder, an action sequence decoder, and an instruction decoder; the visual encoder supports extracting image features from images captured by the drone gimbal camera; the language model supports encoding task instructions and extracting word element embeddings of task instructions as instruction text features; the state encoder supports extracting state features from the drone and gimbal state vectors; The action sequence decoder generates the action sequence of the drone and gimbal based on image features, instruction text features, and state features. The instruction decoder generates the instructions that the drone should execute at the next moment based on image features, instruction text features, and state features.
2. The water conservancy drone inspection method based on the visual language action multimodal model according to claim 1 is characterized in that: The working process of the visual encoder includes: Use the PatchEmbedding method to preprocess the images collected by the UAV gimbal to obtain feature embedding; The visual encoder uses the ViT model to analyze feature embedding and finally output visual features ; Among them, the ViT model is composed of a combination of multi-layer stacked attention and feedforward neural networks; each layer structure is a series of multi-head attention layer, layer normalization operation, feedforward fully connected layer, and layer normalization operation.
3. The water conservancy drone inspection method based on the visual language action multimodal model according to claim 2 is characterized in that: The PatchEmbedding method is used to preprocess the images collected by the UAV gimbal to obtain feature embedding, including: The visual encoder receives images collected by the drone’s gimbal; Divide the image into multiple fixed-size regions. Assume that the region size of each region is (patch_size, patch_size). Extract all regions from the image through a convolution operation with the convolution kernel size as the region size and the step size as patch_size. After the region is divided, a convolutional layer is used to extract the features of each region at the same time. The output of the convolutional layer generates a new tensor representing the regional features. The tensor extracted by the convolutional layer is flattened to merge the spatial dimensions; then, the flattened tensor is transposed; The position encoding is added to the region features to obtain feature embedding.
4. The water conservancy drone inspection method based on the visual language action multimodal model according to claim 3 is characterized in that: The position coding adopts relative position coding: Assume that an input image is divided into regions, for each pair of regions , its relative position encoding is expressed as: ; in, Is used to generate and distance Related encoding functions, This is done using the sine and cosine functions: ; in, is a hyperparameter that controls the frequency of the function. , ; Is a function based on the region position, used to represent the patch and patch The relative position features of : ; in is a learnable weight matrix; All relative position encodings are accumulated to obtain the final output of the position encoding: 。 5. The water conservancy drone inspection method based on the visual language action multimodal model according to claim 1 is characterized in that: The language model adopts the 2B Gemma language model in the PaliGemma VLM model.
6. The water conservancy drone inspection method based on the visual language action multimodal model according to claim 1 is characterized in that: The state encoder maps the input drone and gimbal state vectors to a high-dimensional space through a fully connected layer to increase the expressive power of the features, and then uses The activation function introduces nonlinear transformation, and the final output state characteristics are as follows: ; in, is the input drone and gimbal state vector, 、 is the weight matrix and bias vector of the first layer of the fully connected layer, 、 is the weight matrix and bias vector of the second layer of the fully connected layer, It is the final state characteristic; The drone state vector is represented as: ; in, is the drone state vector, 、 、 is the location of the drone, 、 、 is the component speed of the drone, is the heading angle of the drone, is the drone battery level; The gimbal state vector is expressed as: ; in, is the gimbal state vector, is the pitch angle of the gimbal camera, is the roll angle of the gimbal camera point, is the yaw angle of the gimbal camera, is the zoom factor of the gimbal camera.
7. The water conservancy drone inspection method based on the visual language action multimodal model according to claim 1 is characterized in that: The instruction decoder first processes the visual features , the command text features of task instructions and state characteristics The concatenation is performed, and then the concatenated features are converted into the probability distribution of instructions through the linear layer and Softmax, thereby outputting the instructions that should be executed at the next moment. The calculation process is as follows: ; in, Indicates that the visual features , instruction text features and state characteristics Splicing, is the weight matrix of the linear layer of the instruction decoder, is the bias matrix of the linear layer of the instruction decoder.
8. The water conservancy drone inspection method based on the visual language action multimodal model according to claim 1 is characterized in that: The action sequence decoder is based on the Transformer architecture. The input of the action sequence decoder includes: visual features , instruction text features , state characteristics The action sequence decoder consists of a linear layer, a stack of three multi-head attention layers, a linear layer connected to the attention layer to predict the next action of the drone, and a linear layer connected to the attention layer to predict the next action of the gimbal. When predicting actions, the visual features are maintained. Instruction text features No change, only update the drone status characteristics according to the latest drone status , iterate N times and output the drone actions and gimbal actions at N future moments.
9. The water conservancy drone inspection method based on the visual language action multimodal model according to claim 1 is characterized in that: Based on the basic capabilities required for drone inspections in water conservancy scenarios, basic natural language mission instructions are designed. The mission instructions include: "continuous flight inspection", "hovering", "flying around the target", "taking accurate photos of the target", "flying along the target", and "gimbal tracking the target".
10. A water conservancy drone inspection device based on a visual language action multimodal model, characterized in that: include: At least one processing unit, connecting the processing unit and the storage unit through a bus unit, the storage unit as a computer-readable storage medium, can be used to store software programs, computer executable programs and modules, and the processing unit implements the water conservancy drone inspection method based on the visual language action multimodal model as described in any one of claims 1 to 8 by running the software programs, computer executable programs and modules stored in the storage unit.
Citation Information
Patent Citations
Visual inspection multitask learning method based on multimodal prompt cooperation
CN118918447A
Unmanned aerial vehicle visual language navigation method based on large model task analysis
CN119197530A
Unmanned aerial vehicle visual language navigation method based on visual target reference guidance
CN119245649A
Multi-modal perception problem generation method and system based on large language model, and medium
CN119942300A
Unmanned aerial vehicle multi-modal feature fusion target tracking method and system based on natural language description
CN120013992A
Cited By
Unmanned aerial vehicle resource scheduling method and system based on unmanned aerial vehicle information and multi-modal data
CN121212747A
Multi-unmanned aerial vehicle target coverage method and device based on visual language model
CN122151958A
Multi-uav target coverage method and device based on visual language model
CN122151958B