Unmanned aerial vehicle image water drop removing and optimizing method and device
By optimizing multimodal large model and image recovery algorithm, the image blur problem caused by water droplets attached to the drone's lens is solved, and the intelligent visual perception and task execution capabilities of the drone in bad weather are improved.
Patent Information
- Application Number
- CN202510353474.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-18
AI Technical Summary
In severe weather conditions, the lens attaches water droplets to the lens, causing blurred images, affecting visual perception and decision-making capabilities, and limiting autonomous navigation and obstacle avoidance capabilities.
By collecting and annotating data sets with attached water droplets and clean images, optimizing multimodal large models, designing image recovery algorithms, using information vectors and physical models to remove the influence of water droplets, establishing an agent system, and achieving refined adjustment and generalization capabilities of images.
It improves the visual clarity of the drone under the condition that the lens is attached to water droplets, enhances the task execution capabilities, and provides intelligent decision-making and control flexibility.
Smart Images

Figure CN120339631A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and device for removing and optimizing water droplets in drone images, belonging to the technical field of drones. Background Art
[0002] With the rapid leap of drone technology, its role in numerous fields has become increasingly important. However, the unpredictability of nature, especially under adverse weather conditions, has become a major obstacle to the efficient execution of drone tasks. The image blurring caused by rainwater adhering to the lens seriously affects the visual perception and decision-making ability of drones, posing a severe challenge to complex tasks that require high-precision navigation, obstacle avoidance, and environmental adaptability.
[0003] Traditional drone vision algorithms mainly rely on flight modes with preset GPS routes and compound waypoint control. This mode is difficult to flexibly respond to sudden environmental changes and obstacles, restricting the autonomous navigation and obstacle avoidance capabilities of drones in complex environments. Especially under the harsh visual conditions of raindrops adhering to the lens, due to the lack of perception feedback of the image information of raindrops on the lens, tasks such as fixed-point photography and instrument reading cannot be completed with high quality at the preset waypoints. Moreover, due to the failure to utilize the real-time perception of the surrounding environment by the drone's visual sensor, the drone cannot intelligently plan paths and make reasonable decisions.
[0004] Therefore, developing a technology that can intelligently identify and remove the interference of raindrops in images and combine advanced multi-modal large models for image understanding and effect evaluation has become the key to improving the success rate and reliability of drone tasks. Summary of the Invention
[0005] To solve the above problems, the present invention proposes a method and device for removing and optimizing water droplets in drone images, which can make up for the deficiency of the perception ability of traditional drones when water droplets adhere to the lens during field inspections, realize the intelligent visual perception optimization of drones under the condition of water droplet adhesion, and improve the autonomous navigation, obstacle avoidance ability, and task execution ability of drones.
[0006] The technical solution adopted by the present invention to solve its technical problems is as follows:
[0007] In a first aspect, a method for removing and optimizing water droplets in drone images provided by an embodiment of the present invention includes the following steps:
[0008] Step S1, collect a data set containing images with attached water droplets and clean images, and perform annotation on the data set to construct an instruction fine-tuning data set, where the instruction fine-tuning data set includes discriminative descriptions of water droplets attached to the images and image quality evaluation;
[0009] Step S2, set prompt words, and optimize and train the multimodal large model in combination with the instruction fine-tuning dataset. The trained multimodal large model can not only disassemble the user's instruction into machine instructions and provide top-level task planning information for the drone, but also generate an image description of the input image, where the image description includes whether the image is affected by attached water droplets and the image restoration quality;
[0010] Step S3, input the image into the trained multimodal large model to obtain an image description, use the regular matching method to sequentially extract the key visual elements in the image description, and convert them into information vectors represented digitally;
[0011] Step S4, design an image attached water droplet removal algorithm based on the physical model of attached water droplets and adversarial training to obtain an image restoration model, and use the image restoration model to refine and improve the generalization ability of the image;
[0012] Step S5, optimize the restoration result of the image restoration model based on the information vector of the multimodal large model, and use the optimized image restoration model to remove and optimize the water droplets in the drone image.
[0013] As a possible implementation manner of this embodiment, the step S1 includes the following steps:
[0014] Step S11, determine the overall scale of the dataset, that is, determine the number of images including images with water droplets and images without water droplets;
[0015] Step S12, collect image samples with water droplets and image samples without water droplets, and use the image samples without water droplets as the control group;
[0016] Step S13, perform preprocessing such as scaling and cropping on the collected image samples, and unify the image size and file format;
[0017] Step S14, add annotation information to the image samples with water droplets, where the annotation information includes the position, shape, size and density of the water droplets;
[0018] Step S15, perform image quality evaluation, and assign a quality score to each image according to factors such as clarity, color and contrast;
[0019] Step S16, create an instruction set, write an instruction template for describing image features, and define the structure and syntax of the instruction;
[0020] Step S17, organize the image files and their corresponding label information in the dataset into a structured file to form an instruction fine-tuning dataset;
[0021] Step S18: Validate the instruction fine-tuning dataset to check for errors or missing information in the instruction fine-tuning dataset;
[0022] Step S19: Divide the instruction fine-tuning dataset into a training set, a validation set, and a test set.
[0023] As a possible implementation of this embodiment, step S2 includes the following steps:
[0024] Step S21: Select a pre-trained multi-modal large model as the base model, where the pre-trained multi-modal large model supports the fusion processing of images and texts;
[0025] Step S22: Define the specific tasks that the pre-trained multi-modal large model needs to complete. The specific tasks at least include: disassembling user instructions into machine instructions, providing top-level task planning information for the drone, and generating an image description of the input image;
[0026] Step S23: Use the instruction fine-tuning dataset to prepare the input-output pairs required for training;
[0027] Step S24: Design prompt words to guide the multi-modal large model to generate correct outputs;
[0028] Step S25: Use the training set of the instruction fine-tuning dataset to train the model and fine-tune the hyperparameters of the multi-modal large model;
[0029] Step S26: Use the validation set of the instruction fine-tuning dataset to evaluate the performance of the multi-modal large model;
[0030] Step S27: Use the test set of the instruction fine-tuning dataset to evaluate the multi-modal large model and iterate for improvement.
[0031] As a possible implementation of this embodiment, step S3 includes the following steps:
[0032] Step S31: Input the image into the trained multi-modal large model to obtain an image description, that is, a specific description of the image. The specific description includes whether the image has degradation caused by attached water droplets and whether the image restoration is natural;
[0033] Step S32: Develop a regular expression rule set and combine it with natural language processing algorithms to sequentially extract the key visual elements in the output description of the multi-modal large model;
[0034] Step S33: Convert the extracted key visual elements into digital representations to form an information vector IV. Among them, IV[0] represents whether there are attached raindrops, and IV[1] represents the image restoration quality.
[0035] As a possible implementation manner of this embodiment, step S4 includes the following steps:
[0036] Step S41, analyze the physical model that can characterize the internal mechanism of the degradation of attached raindrops:
[0037] I = J(1 - M)+BM
[0038] Where, I represents the image of the degradation of attached water droplets, J represents the clean image, M is the mask of attached water droplets, which represents the raindrop area when the value is 1, and otherwise represents the clean area; B is the raindrop layer, which represents the complex mixture of the original background and the light reflected by the environment;
[0039] Step S42, derive the image restoration model according to the physical model of the degradation of attached water droplets:
[0040]
[0041] Step S43, design the attached water droplet mask estimation network and the raindrop layer estimation network to obtain the loss function of the image restoration model:
[0042]
[0043] Where, represents the L2 loss, L h,v represents the gradient loss in the horizontal and vertical directions, and λ1 and λ2 respectively represent and L h,v the loss balance coefficients of;
[0044] Step S44, based on the image restoration model, use the parameters estimated by the network to synthesize a preliminary dewatered image;
[0045] Step S45, perform fine-tuning of image restoration based on the encoder-decoder structure to optimize the network end-to-end;
[0046] Step S46, perform adversarial training based on the generative adversarial network, introduce unpaired attached water droplet images and clean images into the attached water droplet degradation restoration training process, and improve the generalization ability of the image restoration algorithm in the real world.
[0047] As a possible implementation manner of this embodiment, step S5 includes the following steps:
[0048] Step S51, store each information vector in a dimension-aligned manner;
[0049] Step S52, use the preliminary restoration result of the image based on the image restoration model as the input and the fine-tuning result as the output to obtain the fine-tuning image restoration model, and use the fine-tuning result as the input and the optimization result as the output to obtain the optimized image restoration model;
[0050] Step S53, use the optimized image restoration model to remove and optimize the water droplets in the UAV images.
[0051] As a possible implementation of this embodiment, the storing each information vector in a dimension-aligned manner includes the following steps:
[0052] Analyze the tail element of the information vector storage queue, that is, the information vector IV at the latest incoming time t t , record the elements at its positions 0 and 1, that is, IV t [0] and IV t [1]; if IV t [0] shows that the image has attached water droplets, it is necessary to start the image restoration module to degenerate the image; if after the degeneration, IV t [1] still shows that the restoration quality is not good, multiple rounds of image optimization are required;
[0053] Set the evolution base IV t [2], record the number of rounds of image restoration optimization, and the evolution base is the element IV at the position 2 of the information vector t [2], take the moving average of the time-series historical data with a window size n of 10, and round down:
[0054]
[0055] As a possible implementation of this embodiment, the method for removing and optimizing the water droplets in the UAV images further includes the following steps:
[0056] Step S6, based on the multi-modal large model and the image restoration model, establish a UAV agent for removing and optimizing the water droplets in the UAV images.
[0057] As a possible implementation of this embodiment, for the method for removing and optimizing the water droplets in the UAV images, step S6 includes the following steps:
[0058] Step S61, use the Nvidia Orin computing card as the computing core and the DJI M350 UAV as the actuator to build a complete hardware system;
[0059] Step S62, the multi-modal large model takes the user instruction and the gimbal image as inputs and the user instruction decomposition and image description as outputs;
[0060] Step S63, take the gimbal image and the information vector as inputs and the restored image as the output to establish a raindrop removal module;
[0061] Step S64, take the information vector and the restored image as inputs and the optimized restored image as the output to establish an image restoration evolution body;
[0062] Step S65: Build and implant functions of visual target detection, visual simultaneous localization and mapping, and path planning. Create running nodes for each module in the Ubuntu system, and complete communication between nodes based on ROS (Robot Operating System), and establish a drone agent.
[0063] In a second aspect, a device for removing and optimizing water droplets in drone images provided by an embodiment of the present invention includes:
[0064] A dataset construction module, configured to collect a dataset including images with attached water droplets and clean images, and perform annotation on the dataset to construct an instruction fine-tuning dataset, where the instruction fine-tuning dataset includes discriminative descriptions of water droplets attached to the images and image quality evaluation;
[0065] A model optimization training module, configured to set prompts, and optimize and train a multi-modal large model in combination with the instruction fine-tuning dataset. The trained multi-modal large model can not only disassemble user instructions into machine instructions and provide top-level task planning information for the drone, but also generate an image description of the input image, where the image description includes whether the image is affected by attached water droplets and the image restoration quality;
[0066] An information vector conversion module, configured to input an image into the trained multi-modal large model to obtain an image description, sequentially extract key visual elements in the image description using a regular matching method, and convert them into information vectors represented digitally;
[0067] A water droplet removal algorithm design module, configured to design an algorithm for removing attached water droplets in an image based on a physical model of attached water droplets and adversarial training to obtain an image restoration model, and use the image restoration model to perform fine-tuning on the image and improve the generalization ability;
[0068] A water droplet removal and optimization module, configured to optimize the restoration result of the image restoration model based on the information vector of the multi-modal large model, and use the optimized image restoration model to perform removal and optimization processing of water droplets in drone images.
[0069] As a possible implementation manner of this embodiment, the device for removing and optimizing water droplets in drone images further includes:
[0070] An agent establishment module, configured to establish a drone agent for removing and optimizing water droplets in drone images based on the multi-modal large model and the image restoration model.
[0071] The beneficial effects of the technical solution of the embodiment of the present invention are as follows:
[0072] By integrating the deep learning-driven image raindrop removal algorithm with the comprehensive analysis ability of the multi-modal large model, the present invention optimizes the visual image quality and ensures visual clarity under the condition of raindrops adhering to the lens. The present invention utilizes the powerful generalization ability, knowledge transfer, and comprehension ability of the multi-modal large model to improve the adaptability of the drone to the environment and its task completion ability. The present invention designs an image restoration evolver with self-iterative and self-memorizing functions, which can optimize the image restoration results based on historical data and continuously improve the restoration quality. The present invention combines multiple modules to construct a complete drone system, realizing the intelligent visual perception optimization of the drone under the condition of raindrops adhering, and improving the autonomous navigation, obstacle avoidance ability, and task execution ability of the drone.
[0073] The present invention ingeniously integrates the image processing mechanism with the control strategy driven by the multi-modal large model and creatively incorporates the large-scale language model to achieve in-depth parsing and precise decomposition of user instructions; by establishing a discriminative description of raindrops adhering to the image and a fine-tuning dataset for image quality evaluation instructions, the system optimizes the generalization ability and understanding ability of the multi-modal large model. Using advanced visual feedback technology, it can efficiently identify the influence of raindrops in the image and generate specific descriptions. At the same time, the present invention can also design an information vector generation module and a raindrop removal algorithm based on the physical model, significantly improving the image restoration quality. The system flexibly optimizes the restoration results through the self-iterative and self-memorizing restoration evolver, establishing an intelligent drone system that can eliminate the influence of raindrops, fundamentally breaking through the limitations of traditional drones when raindrops adhere to the lens in bad weather, and providing intelligent decision-making ability and control flexibility. Brief Description of the Drawings
[0074] Figure 1 is a flowchart of a method for removing and optimizing raindrops in drone images shown according to an exemplary embodiment;
[0075] Figure 2 is a schematic structural diagram of a device for removing and optimizing raindrops in drone images shown according to an exemplary embodiment;
[0076] Figure 3 is a schematic diagram of a system for removing and optimizing raindrops in drone images driven by a multi-modal large model constructed;
[0077] Figure 4 is a schematic diagram showing the degradation of an image when raindrops adhere to the drone lens. Detailed Description of the Embodiment
[0078] To more clearly illustrate the technical features of the solution of the present invention, the present invention will be elaborated in detail below through specific embodiments and in conjunction with its accompanying drawings.
[0079] As Figure 1As shown in the figure, a method for removing and optimizing water droplets in drone images provided by an embodiment of the present invention includes the following steps:
[0080] Step S1, collect a dataset containing images with attached water droplets and clean images, and annotate the dataset to construct an instruction fine-tuning dataset. The instruction fine-tuning dataset includes discriminative descriptions of water droplets attached to the images and image quality evaluations.
[0081] Step S2, set prompt words, and optimize and train a multi-modal large model in combination with the instruction fine-tuning dataset. The trained multi-modal large model can not only disassemble user instructions into machine instructions and provide top-level task planning information for the drone, but also generate an image description of the input image. The image description includes whether the image is affected by attached water droplets and the image restoration quality.
[0082] Step S3, input the image into the trained multi-modal large model to obtain an image description, use the regular matching method to sequentially extract key visual elements in the image description, and convert them into information vectors represented digitally.
[0083] Step S4, design an algorithm for removing attached water droplets in the image based on the physical model of attached water droplets and adversarial training to obtain an image restoration model, and use the image restoration model to refine the adjustment and improve the generalization ability of the image.
[0084] Step S5, optimize the restoration result of the image restoration model based on the information vector of the multi-modal large model, and use the optimized image restoration model to remove and optimize water droplets in the drone image.
[0085] As a possible implementation manner of this embodiment, the step S1 includes the following steps:
[0086] Step S11, determine the overall scale of the dataset, that is, determine the number of images including images with water droplets and images without water droplets.
[0087] Step S12, collect image samples with water droplets and image samples without water droplets, and use the image samples without water droplets as the control group.
[0088] Step S13, perform preprocessing such as scaling and cropping on the collected image samples, and unify the image size and file format.
[0089] Step S14, add annotation information to the image samples with water droplets. The annotation information includes the position, shape, size, and density of the water droplets.
[0090] Step S15, perform image quality evaluation, and assign a quality score to each image according to factors such as clarity, color, and contrast.
[0091] Step S16, create an instruction set, write an instruction template for describing image features, and define the structure and syntax of the instructions;
[0092] Step S17, organize the image files and their corresponding label information in the dataset into a structured file to form an instruction fine-tuning dataset;
[0093] Step S18, verify the instruction fine-tuning dataset to check whether there are errors or missing information in the instruction fine-tuning dataset;
[0094] Step S19, divide the instruction fine-tuning dataset into a training set, a validation set, and a test set.
[0095] As a possible implementation of this embodiment, the step S2 includes the following steps:
[0096] Step S21: Select a pre-trained multi-modal large model as the base model, and the pre-trained multi-modal large model supports the fusion processing of images and texts;
[0097] Step S22: Define the specific tasks that the pre-trained multi-modal large model needs to complete. The specific tasks at least include: disassembling user instructions into machine instructions, providing top-level task planning information for the drone, and generating an image description of the input image;
[0098] Step S23: Use the instruction fine-tuning dataset to prepare the input-output pairs required for training;
[0099] Step S24: Design prompt words to guide the multi-modal large model to generate correct outputs;
[0100] Step S25: Use the training set of the instruction fine-tuning dataset to train the model and fine-tune the hyperparameters of the multi-modal large model;
[0101] Step S26: Use the validation set of the instruction fine-tuning dataset to evaluate the performance of the multi-modal large model;
[0102] Step S27: Use the test set of the instruction fine-tuning dataset to evaluate the multi-modal large model and iterate for improvement.
[0103] As a possible implementation of this embodiment, the step S3 includes the following steps:
[0104] Step S31, input the image into the trained multi-modal large model to obtain a specific description of the image. The specific description includes whether the image has degradation caused by attached water droplets and whether the image restoration is natural;
[0105] Step S32, develop a regular expression rule set, and combine it with natural language processing algorithms to sequentially extract the key visual elements in the output description of the multi-modal large model;
[0106] Step S33: Convert the extracted key visual elements into digital representations to form an information vector IV, where IV[0] represents whether there are attached raindrops, and IV[1] represents the image restoration quality.
[0107] As a possible implementation of this embodiment, step S4 includes the following steps:
[0108] Step S41: Analyze the physical model that can characterize the internal mechanism of attached raindrop degradation:
[0109] I = J(1 - M)+BM
[0110] where I represents the image degraded by attached water droplets, J represents the clean image, M is the attached water droplet mask, which represents the raindrop area when the value is 1, otherwise it is the clean area; B is the raindrop layer, which represents the complex mixture of the original background and the light reflected by the environment.
[0111] Step S42: Derive the image restoration model according to the physical model of attached water droplet degradation:
[0112]
[0113] Step S43: Design an attached water droplet mask estimation network and a raindrop layer estimation network to obtain the loss function of the image restoration model:
[0114]
[0115] where represents the L2 loss, L h,v represents the gradient loss in the horizontal and vertical directions, and λ1 and λ2 respectively represent and L h,v 's loss balance coefficients;
[0116] Step S44: Based on the image restoration model, use the parameters estimated by the network to synthesize a preliminary de-waterdrop image;
[0117] Step S45: Perform fine-tuning of image restoration based on the encoder-decoder structure to optimize the network end-to-end;
[0118] Step S46: Perform adversarial training based on the generative adversarial network, introduce unpaired attached water droplet images and clean images into the attached water droplet degradation restoration training process, and improve the generalization ability of the image restoration algorithm in the real world.
[0119] As a possible implementation of this embodiment, step S5 includes the following steps:
[0120] Step S51: Store each information vector in a dimension-aligned manner;
[0121] In step S52, using the preliminary image restoration result of the image restoration model as the input and the refined result as the output, a refined image restoration model is obtained, and using the refined result as the input and the optimized result as the output, an optimized image restoration model is obtained;
[0122] In step S53, the optimized image restoration model is used to remove and optimize the water droplets in the UAV image.
[0123] As a possible implementation manner of this embodiment, the storing each information vector in a dimension-aligned manner includes the following steps:
[0124] Analyze the tail element of the information vector storage queue, that is, the information vector IV at the latest incoming time t t , and record the elements at its positions 0 and 1, that is, IV t [0] and IV t [1]; if IV t [0] shows that the image has attached water droplets, it is necessary to start the image restoration module to degenerate the image; if after the degeneration, IV t [1] still shows that the restoration quality is not good, it is necessary to perform multiple rounds of image optimization;
[0125] Set the evolution base IV t [2], record the number of rounds of image restoration optimization, and the evolution base is the element IV at the position 2 of the information vector t [2], take the moving average of the time series historical data with a window size n of 10, and round down:
[0126]
[0127] As a possible implementation manner of this embodiment, the method for removing and optimizing the water droplets in the UAV image further includes the following steps:
[0128] In step S6, based on the multimodal large model and the image restoration model, a UAV agent for removing and optimizing the water droplets in the UAV image is established.
[0129] As a possible implementation manner of this embodiment, for the method for removing and optimizing the water droplets in the UAV image, step S6 includes the following steps:
[0130] In step S61, use the Nvidia Orin computing card as the computing core and the DJI M350 UAV as the actuator to build a complete hardware system;
[0131] In step S62, the multimodal large model takes the user instruction and the gimbal image as the input and the user instruction decomposition and the image description as the output;
[0132] Step S63: Taking the gimbal image and the information vector as inputs and the restored image as the output, an attached raindrop removal module is established.
[0133] Step S64: Taking the information vector and the restored image as inputs and the optimized restored image as the output, an image restoration evolver is established.
[0134] Step S65: Construct and implant functions of visual target detection, visual simultaneous localization and mapping, and path planning, create running nodes for each module in the Ubuntu system, and complete communication between nodes based on ROS (Robot Operating System), and establish a drone agent.
[0135] As Figure 2 shown, a device for removing and optimizing water droplets in drone images provided by an embodiment of the present invention includes:
[0136] A dataset construction module for collecting a dataset containing attached water droplet images and clean images, and performing annotation on the dataset to construct an instruction fine-tuning dataset, where the instruction fine-tuning dataset includes discriminative descriptions of image attached water droplets and image quality evaluation.
[0137] A model optimization training module for setting prompts and optimizing and training a multi-modal large model in combination with the instruction fine-tuning dataset. The trained multi-modal large model can not only disassemble user instructions into machine instructions and provide top-level task planning information for the drone, but also generate an image description of the input image, where the image description includes whether the image is affected by attached water droplets and the image restoration quality.
[0138] An information vector conversion module for inputting an image into the trained multi-modal large model to obtain an image description, sequentially extracting key visual elements in the image description using a regular matching method, and converting them into an information vector represented digitally.
[0139] A removal algorithm design module for designing an image attached water droplet removal algorithm based on an attached water droplet physical model and adversarial training to obtain an image restoration model, and using the image restoration model to perform fine-tuning and generalization ability improvement on the image.
[0140] A water droplet removal and optimization module for optimizing the restoration result of the image restoration model based on the information vector of the multi-modal large model, and using the optimized image restoration model to perform water droplet removal and optimization processing on drone images.
[0141] As a possible implementation manner of this embodiment, the device for removing and optimizing water droplets in drone images further includes:
[0142] An agent establishment module, which is used to establish a drone agent for removing and optimizing water droplets in drone images based on a multimodal large model and an image restoration model.
[0143] Figure 3 It is a schematic diagram of a system for removing and optimizing water droplets in drone images driven by a constructed multimodal large model. In these task scenarios where water droplets adhere to the lenses, conventional drone control methods often struggle to meet the actual business requirements and cannot obtain clear and accurate perception information. To solve this problem, the present invention introduces a multimodal large model, an image de-raining algorithm, and an agent construction theory.
[0144] First, an image water droplet adhesion discrimination description and an image quality evaluation instruction fine-tuning dataset are established to provide a data basis for subsequent multimodal large model optimization, thereby fully leveraging the powerful generalization ability, knowledge transfer, and comprehension ability of the large model. The multimodal large model is optimized by setting prompts and an instruction fine-tuning dataset. Through multiple rounds of training and testing, the large model is enabled to have the following capabilities: disassembling user instructions into machine instructions to provide top-level task planning information for the drone; generating specific descriptions of the input image, including whether the image is affected by adhered water droplets and the image restoration quality.
[0145] Then, an information vector generation module is designed. According to the image description output by the multimodal large model, value information in the image description is sequentially extracted using a regular matching method and aggregated into an information vector. An image water droplet removal algorithm based on the physical model of adhered water droplets and adversarial training is designed to improve the restoration quality of the algorithm when facing real scenarios. To meet the requirements in different scenarios, the present invention designs an image restoration evolver with self-iteration and self-memory functions, which can optimize the restoration result based on the information vector of the multimodal large model. Through visual feedback, the drone can perceive changes in the surrounding environment in real time and adjust its behavior according to the environmental changes to better adapt to complex and changing task scenarios.
[0146] Finally, the present invention designs a complete drone image water droplet removal and visual perception optimization system that combines a multimodal large model, an image restoration module based on a physical model, an information vector generation module, an image restoration evolver, and a visual task execution module, and establishes a drone agent driven by a multimodal large model that can eliminate the influence of adhered water droplets.
[0147] Through the above method, the multi-modal large model-driven UAV image water droplet removal and visual perception optimization system implemented by the present invention can significantly improve the intelligence level and task execution efficiency of UAVs when there are water droplets adhering to the lens in the wild scene. This system not only enhances the task planning performance of UAVs by leveraging the powerful capabilities of the multi-modal large model, but also effectively improves the image quality and enhances the visual recognition accuracy when dealing with raindrop-adhering situations. In addition, this system has flexible configuration and expansion capabilities, and can be customized according to the specific needs of different users, providing efficient and accurate solutions for various complex task scenarios.
[0148] In this embodiment, the specific implementation process of using the present invention for UAV image water droplet removal and optimization is as follows.
[0149] S1: Establish a discriminative description of water droplets adhering to the image and a fine-tuning data set for image quality evaluation instructions, providing a data basis for subsequent optimization of the multi-modal large model, so as to fully utilize the powerful generalization ability, knowledge transfer, and comprehension ability of the large model.
[0150] Step S1 includes the following detailed steps:
[0151] S11: Determine the overall scale of the data set, that is, determine the number of water droplet-bearing images and non-water droplet-bearing images required;
[0152] S12: Collect image samples with water droplets and collect non-water droplet-bearing image samples as a control group; ensure that the image sources are diverse, including photos in different environments, such as outdoors, indoors, and different weather conditions, etc.;
[0153] S13: Perform necessary preprocessing on the images, including scaling, cropping, etc., to unify the image size and ensure that the image file formats are consistent;
[0154] S14: Manually annotate the positions of water droplets in each water droplet-bearing image using an annotation tool, and add information such as the shape, size, and density of the annotated water droplets;
[0155] S15: Conduct image quality evaluation, assign a quality score to each image, considering factors such as clarity, color, and contrast, and an expert database or inviting experts can be used for quality scoring;
[0156] S16: Create an instruction set, write an instruction template for describing image features, such as: "The image clarity is good", "The image is severely affected by water droplets", etc., and define the structure and syntax of the instructions;
[0157] S17: Organize the data set, organize the image files and their corresponding label information into a structured file, such as a CSV or JSON file, to ensure that each image has a clear identifier and corresponding label information;
[0158] S18: Dataset verification, check for errors or missing information in the dataset; use a small portion of the data for preliminary testing to ensure the dataset can be used for training the model;
[0159] S19: Dataset partitioning, partition the dataset into a training set, a validation set, and a test set, with proportions of 70%, 15%, and 15%.
[0160] S2: Optimize the multi-modal large model by setting prompt words and instructions to fine-tune the dataset, and through multiple rounds of training and testing, enable the large model to have the following capabilities: disassemble user instructions into machine instructions and provide top-level task planning information for the drone; generate specific descriptions of the input image, including whether the image is affected by attached water droplets and the image restoration quality.
[0161] Step S2 includes the following detailed steps:
[0162] S21: Select a pre-trained multi-modal large model as the base model and confirm that the model supports the fusion processing of images and text;
[0163] S22: Define the specific tasks that the model needs to complete, such as: disassemble user instructions into machine instructions; provide top-level task planning information for the drone; generate specific descriptions of the input image, including the presence or absence of water droplets and the image restoration quality;
[0164] S23: Use the dataset from the S1 stage to prepare the input-output pairs required for training and ensure that the dataset has been preprocessed as required;
[0165] S24: Design prompt words for guiding the model to generate correct outputs, and the prompt words should be able to help the model understand the user's intention and generate corresponding descriptions;
[0166] S25: Model fine-tuning, start fine-tuning the model using the training dataset and adjust the hyperparameters of the model as needed, such as the learning rate, batch size, etc.;
[0167] S26: Model evaluation, evaluate the performance of the model on the validation set, and use metrics such as accuracy, recall, and F1-score to measure the performance of the model;
[0168] S27: Model testing, conduct a final evaluation of the model on an independent test set to ensure that the model can generalize to unseen data;
[0169] S28: Iterative improvement, based on the performance of the model on the test set, identify deficiencies; adjust the model architecture or training strategy, retrain and evaluate the model;
[0170] S29: Deploy the model. After the model meets the performance requirements, deploy the model for actual use and set up a monitoring mechanism to continuously evaluate the model's performance in the actual environment.
[0171] S3: Design an information vector generation module. According to the image descriptions output by the multimodal large model, use the regular matching method to sequentially extract the valuable information in the image descriptions and aggregate it into an information vector.
[0172] Step S3 includes the following detailed steps:
[0173] S31: Input the image into the multimodal large model fine-tuned with instructions to obtain a specific description of the image, including: whether the image has degradation caused by attached water droplets, that is, to determine whether the lens has water droplets; whether the image restoration is natural, that is, the image quality; as Figure 4 shown, an example of the image description is as follows: "The image is degraded due to attached water droplets, and the image quality is poor";
[0174] S32: Develop a regular expression rule set, combined with natural language processing algorithms, to sequentially extract the key visual elements in the descriptions output by the multimodal large model to accurately capture the words and phrases directly related to the image restoration task;
[0175] S33: Convert the extracted valuable information into a digital representation to form an information vector IV. Among them, IV[0] represents whether there are attached raindrops, with a value of 1 indicating that there are attached raindrops and 0 indicating that the image is clean; IV[1] represents the image restoration quality, with a value of 1 indicating high restoration quality and 0 indicating that further restoration is required.
[0176] S4: Design an image attached water droplet removal algorithm based on the physical model of attached water droplets and adversarial training to improve the restoration quality of the algorithm when facing real scenarios.
[0177] Step S4 includes the following detailed steps:
[0178] S41: Analyze the physical model that can characterize the internal mechanism of attached raindrop degradation, as shown in Equation (1):
[0179] I = J(1 - M) + BM (1)
[0180] Where, I represents the image degraded by attached water droplets, J represents the clean image, M is the attached water droplet mask, with a value of 1 indicating the raindrop area and otherwise the clean area; B is the raindrop layer, representing the complex mixture of the original background and the light reflected by the environment.
[0181] S42: Derive an image restoration model according to the physical model of attached water droplet degradation, as shown in Equation (2):
[0182]
[0183] S43: Design the attached water droplet mask estimation network and the raindrop layer estimation network, namely M-Net and B-Net, to provide basic parameters for synthesizing and restoring images. The network structure adopts a 12-layer Transformer architecture. Taking the parameter B as an example, the loss function is shown in Equation (3):
[0184]
[0185] Among them, represents the L2 loss, and L h,v represents the gradient loss in the horizontal and vertical directions. λ1 and λ2 respectively represent and L h,v 's loss balance coefficients; m and n represent the image size, B and B gt represent the predicted value and the true value of the raindrop layer, D x and D y represent the gradient operators in the horizontal and vertical directions respectively;
[0186] S44: Based on the image restoration model, use the M and B parameters estimated by the network to synthesize a preliminary de-waterdrop image;
[0187] S45: Design an image restoration refinement adjustment module based on the encoder-decoder structure, and use the synthesized real degradation-clean image pair to optimize the network end-to-end, realize the supervision of the image restoration process, and complete the refinement of the restored image;
[0188] S46: Design an adversarial training module based on the generative adversarial network, introduce unpaired attached water droplet images and clean images into the attached water droplet degradation restoration training process, and improve the generalization ability of the image restoration algorithm in the real world.
[0189] S5: Design an image restoration evolver with self-iteration and self-memory functions, which can optimize the restoration result based on the information vector of the multi-modal large model.
[0190] Step S5 includes the following detailed steps:
[0191] S51: Set up an information vector memory, store each information vector in a dimension-aligned manner, and the memory queue length does not exceed 10; this memory has the following functions: analyze the element at the end of the queue, that is, the information vector IV t at the latest incoming time t, and record the elements at its positions 0 and 1, that is, IV t [0] and IV t [1]. If IV t [0] shows that the image has attached water droplets, it is necessary to start the image restoration module to degenerate the image; if after degeneration, IV t[1] If the restored quality is still shown to be poor, then multiple rounds of image optimization are required; in addition, set the evolution base IV t [2], record the number of rounds of image restoration optimization, which is used to provide a reference for subsequent image optimization;
[0192] The evolution base is the element at the position of the information vector 2, i.e., IV t [2], and its calculation is shown in Equation (6). Take the moving average of the time-series historical data with a window size n of 10 and round down:
[0193]
[0194] S52: The refinement adjustment module takes the preliminary image restoration result based on the physical model as input and the refined result as output. The adversarial training module takes the refined result and the optimized result as output, that is, the two modules are connected in series;
[0195] S53: Integrate the information vector memory, the refinement adjustment module, and the adversarial training module to use IV t [0] to control whether to start image restoration, use IV t [1] to control whether multiple rounds of optimization are required, use IV t [2] to provide the evolution base, that is, the historical reference of the number of optimization rounds, so that the image restoration gradually evolves based on the historical data, forming an image restoration evolution body with self-iteration and self-memory functions.
[0196] S6: Design a complete UAV image water droplet removal and visual perception optimization system that combines a multi-modal large model, an image restoration module based on a physical model, an information vector generation module, an image restoration evolution body, and a visual task execution module, and establish a UAV intelligent agent driven by a multi-modal large model that can eliminate the influence of attached water droplets.
[0197] Step S6 includes the following detailed steps:
[0198] S61: Use the Nvidia Orin computing card as the computing core and the DJI drone M350 as the actuator to build a complete hardware system;
[0199] S62: The multi-modal large model takes the user instruction and the gimbal image as input and the user instruction decomposition and image description as output;
[0200] S63: The attached raindrop removal module takes the gimbal image and the information vector as input and the restored image as output;
[0201] S64: The image restoration evolution body takes the information vector and the restored image as input and the optimized restored image as output;
[0202] S65: Build and insert functional modules such as visual object detection, visual simultaneous localization and mapping, and path planning, enabling the drone to have the ability to execute tasks;
[0203] S65: Create running nodes for each module in the Ubuntu system and complete the communication between nodes based on ROS to realize the establishment of the drone agent system.
[0204] The present invention ingeniously integrates the image processing mechanism and the control strategy driven by the multi-modal large model, and creatively incorporates the large-scale language model to achieve in-depth parsing and precise decomposition of user instructions; by establishing the discriminative description of water droplets attached to the image and the instruction fine-tuning data set for image quality evaluation, the system optimizes the generalization ability and understanding ability of the multi-modal large model. Using advanced visual feedback technology, it can efficiently identify the influence of water droplets in the image and generate specific descriptions. At the same time, the present invention can also design an information vector generation module and a water droplet removal algorithm based on the physical model, significantly improving the image restoration quality. The system flexibly optimizes the restoration result through the self-iterative and self-memorizing restoration evolution body, and establishes an intelligent drone system that can eliminate the influence of water droplets, fundamentally breaking through the limitations of traditional drones when water droplets are attached to the lens in bad weather, and providing intelligent decision-making ability and control flexibility.
[0205] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that it is still possible to modify the specific implementation manners of the present invention or make equivalent replacements. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. A method for removing and optimizing water droplets in UAV images, characterized in that, It includes the following steps: Step S1, collect a dataset containing images with attached water droplets and clean images, and annotate the dataset to construct an instruction fine-tuning dataset. The instruction fine-tuning dataset includes discriminative descriptions of water droplets attached to the images and image quality evaluations; Step S2, set prompt words, and optimize and train the multi-modal large model in combination with the instruction fine-tuning dataset. The trained multi-modal large model can disassemble user instructions into machine instructions, provide top-level task planning information for the drone, and generate an image description of the input image. The image description includes whether the image is affected by attached water droplets and the image restoration quality; Step S3, input the image into the trained multi-modal large model to obtain an image description, use the regular matching method to sequentially extract key visual elements in the image description, and convert them into information vectors represented digitally; Step S4, design an image water droplet removal algorithm based on the physical model of attached water droplets and adversarial training to obtain an image restoration model, and use the image restoration model to refine the adjustment and improve the generalization ability of the image; Step S5, optimize the restoration result of the image restoration model based on the information vector of the multi-modal large model, and use the optimized image restoration model to remove and optimize the water droplets in the drone image.
2. The method for removing and optimizing water droplets in UAV images according to claim 1, wherein, The step S1 includes the following steps: Step S11, determine the overall scale of the dataset, that is, determine the number of images including images with water droplets and images without water droplets; Step S12, collect image samples with water droplets and image samples without water droplets, and use the image samples without water droplets as the control group; Step S13, perform preprocessing such as scaling and cropping on the collected image samples, and unify the image size and file format; Step S14, add annotation information to the image samples with water droplets. The annotation information includes the position, shape, size, and density of the water droplets; Step S15, perform image quality evaluation, and assign a quality score to each image according to factors such as clarity, color, and contrast; Step S16, create an instruction set, write an instruction template for describing image features, and define the structure and syntax of the instructions; Step S17, organize the image files and their corresponding label information in the dataset into a structured file to form an instruction fine-tuning dataset; Step S18, verify the instruction fine-tuning dataset to check whether there is incorrect or missing information in the instruction fine-tuning dataset; Step S19, divide the instruction fine-tuning dataset into a training set, a validation set, and a test set.
3. The method for removing and optimizing water droplets in drone images according to claim 1, characterized in that, The step S2 includes the following steps: Step S21: Select a pre-trained multi-modal large model as the basic model. The pre-trained multi-modal large model supports the fusion processing of images and texts; Step S22: Clarify the specific tasks that the pre-trained multi-modal large model needs to complete. The specific tasks include at least: disassembling user instructions into machine instructions, providing top-level task planning information for the drone, and generating an image description of the input image; Step S23: Use the instruction fine-tuning dataset to prepare the input-output pairs required for training; Step S24: Design prompt words to guide the multi-modal large model to generate correct outputs; Step S25: Use the training set of the instruction fine-tuning dataset to train the model and fine-tune the hyperparameters of the multi-modal large model; Step S26: Use the validation set of the instruction fine-tuning dataset to evaluate the performance of the multi-modal large model; Step S27: Use the test set of the instruction fine-tuning dataset to evaluate the multi-modal large model and iterate for improvement.
4. The method for removing and optimizing water droplets in drone images according to claim 1, wherein, The said step S3 includes the following steps: Step S31, input the image into the trained multi-modal large model to obtain an image description, where the image description includes whether the image has degradation caused by attached water droplets and whether the image restoration is natural; Step S32, develop a regular expression rule set, and combine with natural language processing algorithms to sequentially extract key visual elements in the output description of the multi-modal large model; Step S33, convert the extracted key visual elements into digital representations to form an information vector IV, where IV[0] represents whether there are attached raindrops, and IV[1] represents the image restoration quality.
5. The method for removing and optimizing water droplets in drone images according to claim 1, wherein, The said step S4 includes the following steps: Step S41, analyze the physical model that can characterize the internal mechanism of attached raindrop degradation: I = J(1 - M)+BM where I represents the attached water droplet degraded image, J represents the clean image, M is the attached water droplet mask, when the value is 1, it represents the raindrop area, otherwise it is the clean area; B is the raindrop layer, representing the complex mixture of the original background and the light reflected by the environment; Step S42, derive an image restoration model based on the attached water droplet degradation physical model; Step S43, design an attached water droplet mask estimation network and a raindrop layer estimation network to obtain the loss function of the image restoration model; Among them, represents the L2 loss, and L h,v represents the gradient loss in the horizontal and vertical directions. λ1 and λ2 respectively represent and L h,v 's loss balance coefficients; Step S44, based on the image restoration model, use the parameters estimated by the network to synthesize a preliminary de-waterdrop image; Step S45, perform refined adjustment of image restoration based on the encoder-decoder structure and optimize the network end-to-end; Step S46, perform adversarial training based on the generative adversarial network, introduce unpaired attached water droplet images and clean images into the attached water droplet degradation restoration training process, and improve the generalization ability of the image restoration algorithm in the real world.
6. The method for removing and optimizing water droplets in UAV images according to claim 1, characterized in that The said step S5 includes the following steps: Step S51, store each information vector in a dimension-aligned manner; Step S52, use the preliminary image restoration result based on the image restoration model as the input and the refined result as the output to obtain a refined image restoration model, and use the refined result as the input and the optimized result as the output to obtain an optimized image restoration model; Step S53, use the optimized image restoration model to remove and optimize the water droplets in the drone image.
7. The method for removing and optimizing water droplets in drone images according to any one of claims 1-6, characterized in that, It also includes the following steps: Step S6, based on the multi-modal large model and the image restoration model, establish a drone intelligent agent for removing and optimizing water droplets in drone images.
8. The method for removing and optimizing water droplets in UAV images according to claim 7, characterized in that, The said step S6 includes the following steps: Step S61, use the Nvidia Orin computing card as the computing core and the DJI drone M350 as the actuator to build a complete hardware system; Step S62, the multi-modal large model takes the user instruction and the gimbal image as the input and outputs the user instruction decomposition and the image description; Step S63: Establish an attached raindrop removal module with the gimbal image and the information vector as the input and the restored image as the output. Step S64: Establish an image restoration evolution body with the information vector and the restored image as the input and the optimized restored image as the output. Step S65: Construct and implant functions for visual target detection, visual simultaneous localization and mapping, and path planning. Create running nodes for each module in the Ubuntu system and complete communication between nodes based on ROS to establish a drone agent.
9. An apparatus for removing and optimizing water droplets in UAV images, characterized in that, It includes: A dataset construction module for collecting a dataset containing images with attached water droplets and clean images, and annotating the dataset to construct an instruction fine-tuning dataset. The instruction fine-tuning dataset includes discriminative descriptions of image attached water droplets and image quality evaluation. A model optimization training module for setting prompts and optimizing and training a multimodal large model in combination with the instruction fine-tuning dataset. The trained multimodal large model can disassemble user instructions into machine instructions, provide top-level task planning information for the drone, and generate an image description of the input image. The image description includes whether the image is affected by attached water droplets and the image restoration quality. An information vector conversion module for inputting an image into the trained multimodal large model to obtain an image description, sequentially extracting key visual elements in the image description using a regular matching method, and converting them into an information vector represented digitally. A removal algorithm design module for designing an image attached water droplet removal algorithm based on the physical model of attached water droplets and adversarial training to obtain an image restoration model, and using the image restoration model to refine and improve the generalization ability of the image. A water droplet removal and optimization module for optimizing the restoration result of the image restoration model based on the information vector of the multimodal large model, and using the optimized image restoration model to remove and optimize the water droplets in the drone image.
10. The device for removing and optimizing water droplets in drone images according to claim 9, wherein, It also includes: An agent establishment module for establishing a drone agent for removing and optimizing water droplets in drone images based on the multimodal large model and the image restoration model.