Multi-modal image sample generation method and device based on view angle of unmanned aerial vehicle, and medium

By combining UAV flight parameters and a multimodal large model, multimodal image samples adapted to the UAV's perspective are generated, solving the problems of difficulty in collecting small targets and data scarcity in UAV inspections, improving the model's recognition accuracy and generalization ability in complex scenarios, and reducing data collection costs.

CN120894645APending Publication Date: 2025-11-04CASCO SIGNAL LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510973972.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Existing drone inspection technologies face challenges such as difficulty in acquiring images of small targets and a lack of data in zero-sample or small-sample scenarios, resulting in poor target detection models in complex environments and high costs associated with manual data collection.

Method used

By combining UAV flight parameters and multimodal large models, ground sampling distance (GSD) is calculated to generate multimodal image samples, including target segmentation, scaling and multi-background synthesis, to adapt to complex scenes from the UAV's perspective. Diffusion Model and Segment Anything Model (SAM) are used to generate a variety of appearance variations, forming diverse image samples.

Benefits of technology

It improves the generalization ability and detection accuracy of the UAV inspection target detection model in complex scenarios, reduces data collection costs, adapts to zero-sample and small-sample scenarios, and generates samples with high matching degree with actual scenarios, thereby improving the model's recognition effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120894645A_ABST
    Figure CN120894645A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal image sample generation method and device based on a visual angle of an unmanned aerial vehicle, and a medium. The method comprises the following steps: determining a target object polled by the unmanned aerial vehicle and obtaining information of the target object; setting flight parameters of the unmanned aerial vehicle for executing the inspection task, calculating a ground sampling distance GSD based on the flight parameters, and determining an actual ground size corresponding to each pixel in an image shot by the unmanned aerial vehicle; aiming at the condition that an existing sample image exists and the condition that no sample image exists, generating a sample picture of the target object through a multi-modal large model; performing image segmentation processing on the sample picture for generating the target object to generate a segmentation mask, and completely separating the target object from the original background; and scaling the separated target object based on the GSD and the information of the target object, and synthesizing the scaled target object with a multi-background image to generate a multi-modal image sample conforming to a real inspection scene. Compared with the prior art, the method has the advantages of improving data acquisition efficiency and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of UAV inspection image generation, and in particular to a method, device and medium for generating multimodal image samples based on the perspective of a UAV. Background Technology

[0002] In the field of UAV inspection, the acquisition and generation of high-quality image samples is a key step in improving the performance of target detection models. Existing technologies face two major challenges in generating UAV inspection image samples: First, the shooting angle and imaging characteristics of UAVs flying at high altitudes determine that targets occupy a very small proportion in the image, usually appearing as small targets. Traditional sample generation methods, however, often focus on large targets, making it difficult to accurately simulate the shooting angle and imaging characteristics under specific flight conditions. This results in deviations between the generated samples and the actual inspection scenarios, affecting the model training effect. Second, the problem of data scarcity is significant in zero-sample and small-sample scenarios. Traditional solutions that rely on large amounts of real-world data collection are not only costly but also difficult to cover diverse inspection environments such as complex weather and terrain.

[0003] A search revealed Chinese Patent Publication No. CN112990335A, which discloses a self-learning training method and system for intelligent image recognition of power grid drone inspections. This method utilizes defect images captured by a mobile phone-operated drone during inspections. The defect images are then filtered and labeled to establish a defect sample library. Samples from this library are extracted to generate a dataset, and an algorithm model is trained on this dataset to generate a recognition model.

[0004] Chinese patent number CN113240767A discloses a method and system for generating abnormal crack samples in power transmission line channels. By simulating a cloud-to-ground lightning model, it solves the problem that the small number of crack datasets cannot meet the sample quantity requirements of deep learning crack detection, thereby improving the model recognition effect.

[0005] Chinese Patent No. CN112967248A discloses a method for generating defect image samples, comprising: acquiring a target image and location labels corresponding to defective portions of insulators in the target image; determining the matching degree between image blocks in the target image and image blocks in a preset tile library, wherein the tile library is constructed from image blocks acquired in a first defect image sample; constructing a mask matrix based on the matching degree and a preset matching degree threshold; determining a background-free image corresponding to an insulator image in the target image based on the mask matrix and location labels; and generating a second defect image sample based on the background-free image corresponding to the insulator image and a preset normal image sample.

[0006] Chinese patent CN114693665A discloses an insulator defect detection method and system. By acquiring multiple real image samples of insulator defects, artificial samples of insulator defects with different backgrounds and shapes are constructed. The real human samples and artificial image samples are used as inputs to train an insulator defect detection model.

[0007] In summary, existing drone detection technologies and methods rely on the quantity and quality of image samples, and their core value lies in how to construct high-quality sample data. For computer vision, the more similar the shooting perspectives of the training data and the recognition data, the better the recognition effect. Generally, drone image data collected from low-flying perspectives is not suitable as a sample dataset for high-altitude inspection. Even adding multi-scale variation effects to the recognition algorithm cannot eliminate the problem of inaccurate recognition results caused by changes in flight conditions. In addition, existing technologies rely on existing defective samples and cannot address the needs of zero-sample or small-sample scenarios.

[0008] Therefore, how to significantly expand the data sample size at a low cost in scenarios with small or even zero samples, generate multimodal image samples that accurately match the perspective of drone inspection, and thus improve the drone inspection and recognition effect, has become an urgent technical problem to be solved. Summary of the Invention

[0009] The purpose of this invention is to overcome the shortcomings of the existing technology and provide a method, device and medium for generating multimodal image samples based on the perspective of a UAV.

[0010] The objective of this invention can be achieved through the following technical solutions:

[0011] According to a first aspect of the present invention, a method for generating multimodal image samples based on a UAV perspective is provided, the method comprising:

[0012] Identify the drone inspection target and obtain its appearance features and size specifications;

[0013] Determine the flight parameters for the UAV to perform the inspection mission, calculate the ground sampling distance GSD based on the flight parameters, and determine the actual ground size corresponding to each pixel in the image captured by the UAV.

[0014] For cases where existing sample images exist and cases where no sample images exist, multiple appearance variations of the inspection target are generated using a multimodal large model.

[0015] The generated inspection target is subjected to image segmentation processing to generate a segmentation mask, which completely separates the inspection target from the original background;

[0016] Based on the size and specifications of the GSD and the inspection target, the target area is scaled proportionally and synthesized with multiple background images to generate a multimodal image sample that conforms to the real inspection scenario.

[0017] As a preferred technical solution, the flight parameters for the UAV to perform inspection tasks include flight altitude, shooting angle, flight speed, and sensor information carried by the UAV.

[0018] As a preferred technical solution, the sensor information carried by the drone includes image pixel L. p ×W p a) Pixel size information, lens focal length, and lens type.

[0019] As a preferred technical solution, the formula for calculating the ground sampling distance (GSD) is as follows:

[0020]

[0021] Where f is the focal length of the camera lens, a is the pixel size, and H is the flight altitude.

[0022] As a preferred technical solution, the various appearance variations of the inspection target include:

[0023] When an existing sample image exists, the sample image is cropped to ensure that the inspection target occupies more than 50% of the entire image.

[0024] When no existing sample images exist, the inspection targets of the drone are generated using a multimodal large model based on the samples provided from the drone's perspective.

[0025] A multimodal large model based on Diffusion Model is adopted. Through a conditional guidance mechanism, the color and texture of the truncated inspection target and the inspection target generated by the multimodal large model are processed to generate multiple appearance variants of the inspection target.

[0026] As a preferred technical solution, the image segmentation process uses Segment Anything Model (SAM) to automatically identify the target region in the image and generate an accurate segmentation mask, completely separating the inspection target from the original background and outputting independent target material for subsequent background compositing.

[0027] As a preferred technical solution, the scaling is specifically as follows:

[0028] Based on the actual size and GSD value of the inspected target, calculate the target's proportion in the entire image, and scale the target area accordingly:

[0029]

[0030] Where L p W represents the number of horizontal pixels in the drone's image. p L represents the number of vertical pixels in the drone's image. obj W represents the true length of the target. obj This represents the actual width of the target.

[0031] As a preferred technical solution, the multiple background images include:

[0032] Images collected from actual inspection scenarios, public datasets, or computer-generated virtual scenes, covering a variety of weather conditions, terrain, and light intensity.

[0033] As a preferred technical solution, the method for generating multimodal image samples based on UAV perspective further includes the following steps:

[0034] The generated image samples are divided into training set, validation set and test set;

[0035] Training and validation were performed using the YOLO object detection model.

[0036] As a preferred technical solution, the method for generating multimodal image samples based on UAV perspective, after model verification, deploys the trained model to the UAV inspection system to perform real-time target detection.

[0037] According to a second aspect of the present invention, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the program to implement the method described thereon.

[0038] According to a third aspect of the present invention, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method described thereon.

[0039] Compared with the prior art, the present invention has the following advantages:

[0040] (1) This invention combines UAV flight information, airborne sensor information and multimodal large model image generation technology to calculate GSD accuracy and simulate sample images under different flight conditions of UAV, thereby replacing manual collection, improving data collection efficiency and shortening the engineering cycle.

[0041] (2) This invention proposes a method for generating UAV inspection samples that is adapted to zero-sample and small-sample scenarios. It can simultaneously accommodate new inspection scenarios without original samples and special target detection needs with very small sample sizes. By generating multiple appearance variations of inspection targets through a multimodal large model, it eliminates the need for manual collection of a large amount of basic data, avoids additional time and manpower costs, and effectively improves the efficiency of inspection data preparation.

[0042] (3) By combining UAV flight parameters with GSD accuracy calculation, this invention automatically completes the scaling and viewing angle adaptation of the inspection target, thereby improving the matching degree between the sample and the actual scene.

[0043] (4) The sample images generated by this invention are diverse, which can significantly improve the generalization ability and detection accuracy of the target detection model in complex scenes.

[0044] (5) This invention generates multiple appearance variations of inspection targets and synthesizes multiple backgrounds by multimodal large model, which can adapt to the diverse target requirements in complex inspection scenarios. At the same time, it is compatible with sample generation under different flight conditions. The generated diverse samples can significantly improve the generalization ability and detection accuracy of the target detection model and have a certain degree of universality. Attached Figure Description

[0045] Figure 1 This is a schematic diagram of the method flow of the present invention;

[0046] Figure 2 This is a schematic diagram illustrating the generation of color variants from sample images based on a multimodal large model, as per the present invention.

[0047] Figure 3 This is a sample image of the UAV inspection top-down view generated based on a multimodal large model according to the present invention.

[0048] Figure 4 This is a sample image of an unmanned aerial vehicle (UAV) inspection from a top-down perspective, generated based on a multimodal large model according to the present invention, which contains light and shadow information.

[0049] Figure 5 This is a schematic diagram of the target segmentation results of the SAM large model of the present invention;

[0050] Figure 6 This is an example image of a sample image synthesized from multiple backgrounds according to the present invention. Detailed Implementation

[0051] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0052] This invention provides a method for generating multimodal image samples based on the perspective of a drone. In drone inspection scenarios with zero or small samples, it combines a multimodal large model to generate diverse appearance variations of the inspection target, calculates the GSD through flight parameters, adjusts the target scale, and synthesizes multiple backgrounds to generate diverse samples adapted to the drone's perspective.

[0053] like Figure 1 As shown, the specific process of this invention includes the following steps:

[0054] Step S1: Determine the drone inspection target and clarify the specific target objects to be inspected by the drone;

[0055] Step S2: Determine the UAV flight conditions and GSD accuracy;

[0056] Step S3: Generate a small number of target samples based on the multimodal large model;

[0057] Step S4: Obtain the segmentation mask of the detected target sample based on the SAM large model;

[0058] Step S5: Based on GSD precision scaling and multi-background compositing;

[0059] Step S6: Training and validation based on the YOLO model.

[0060] Step S1 specifically includes:

[0061] Step S1-1: Define the specific target objects for drone inspection, such as excavators, power equipment, road facilities, etc.

[0062] Steps S1-2 record the basic information such as the appearance features and size specifications of the target object to provide a benchmark for subsequent sample generation.

[0063] Step S2 specifically includes:

[0064] Step S2-1: Based on the actual inspection scenario, set the key flight parameters for the UAV to perform the inspection task, including flight altitude, shooting angle, and flight speed.

[0065] Step S2-2: Record the sensor information carried by the aircraft, including image pixel L. p ×W p Pixel size information (a), lens focal length, lens type, etc.

[0066] Steps S2-3: The drone's flight parameters directly determine the viewing angle, resolution, and imaging range of the images captured by the drone, forming the basis for simulating real-world inspection scenarios. Based on the drone's flight conditions, the Ground Sampling Distance (GSD) is calculated. The GSD calculation formula is:

[0067]

[0068] Where f is the focal length of the camera lens, a is the pixel size, and H is the flight altitude. By accurately calculating the GSD, the actual ground size corresponding to each pixel in the image can be determined, providing a precise proportional basis for subsequent sample scaling.

[0069] Step S3 specifically includes:

[0070] Step S3-1: Based on the samples provided from the UAV's perspective, use a multimodal large model to generate the UAV's inspection targets.

[0071] Step S3-2: If there are already some sample images of the corresponding viewpoint, the sample images need to be cropped to ensure that the target sample accounts for more than 50% of the entire image.

[0072] Step S3-3: For the generated target or the original sample, perform color and texture transformation processing using a multimodal large model based on Diffusion Model to generate sample images. Multiple appearance variations of the target object are generated through a conditional guidance mechanism.

[0073] Step S4 uses the Segment Anything Model to segment the target object in the color-transformed image. The SAM large model can automatically identify the target region in the image and generate an accurate segmentation mask, completely separating the target object from the original background, providing independent target material for subsequent background replacement and sample synthesis.

[0074] Step S5 specifically includes:

[0075] Step S5-1: Based on the GSD accuracy calculated in step S2-2, the target size information recorded in step S1-2, and the sensor information recorded in S2-2, the proportion of the target in the entire image can be obtained:

[0076]

[0077] Where L p W represents the number of horizontal pixels in the drone's image. p L represents the number of vertical pixels in the drone's image. obj W represents the true length of the target. obj This represents the actual width of the target.

[0078] Step S5-2: Then, based on the target's proportion in the entire image, the target and background are merged to synthesize a new sample image. The background image can be derived from images collected in actual inspection scenarios, public datasets, or computer-generated virtual scenes, covering various weather conditions such as sunny, cloudy, and rainy days, as well as environmental factors such as different terrains and light intensities. By adjusting the scaling ratio of the target object, its size in the synthesized image is ensured to be consistent with the actual target size captured by the drone under set flight conditions, thereby generating multimodal image samples that conform to real inspection scenarios.

[0079] Step S6 specifically includes:

[0080] Step S6-1: Divide the dataset generated in step S5 into three parts: 70% as the training set, 15% as the validation set, and the remaining 15% as the test set.

[0081] Step S6-2: Select the YOLO model for training;

[0082] Step S6-2: Fix the parameters and train the model;

[0083] Step S6-3: Evaluate the model on the validation set;

[0084] Step S6-4, Model Inference and Deployment.

[0085] This invention enables the automatic generation of multimodal samples adapted to different flight conditions in UAV inspection scenarios. By combining a large multimodal model to generate variant samples, using flight parameters to calculate GSD to adjust the target scale, and synthesizing multiple backgrounds, it eliminates the need for a large amount of manually collected data. It can adapt to new inspection scenarios with zero samples and special target detection scenarios with small samples, significantly improving the practicality of the solution in complex inspection scenarios. Specific Implementation

[0087] by Figures 2-6 For example, the steps of the present invention will be explained in detail below:

[0088] This is achieved based on the following conditions:

[0089] The inspection target was the excavator;

[0090] The drone used was a DJI Mitrice 3TD, with each image being 48 megapixels. The drone was set to fly at an altitude of 100 meters, with a 90-degree overhead shooting angle. It used a 3x zoom lens and a full-frame camera (focal length f=24mm) with a 1 / 1.3-inch CMOS sensor, estimated to be 9.6mm × 7.2mm in size, with an actual diagonal of approximately 12.2mm.

[0091] The specific steps are as follows:

[0092] The target of the drone inspection was identified as an excavator, and the approximate length, width, and other dimensions of the excavator were recorded.

[0093] Determine the flight parameters for the UAV to perform the inspection mission:

[0094] The flight altitude H is 100 meters, and the shooting angle is a 90-degree overhead view;

[0095] By analyzing an image with 48 megapixels, the image pixel count L can be derived. p ×W p It is 8000×6000;

[0096] Pixel size is estimated based on sensor size.

[0097]

[0098] When the lens uses three times the focal length, the actual focal length of the sensor is f = 6.8mm × 3 = 20.4mm;

[0099] Calculate GSD accuracy:

[0100]

[0101] That is, 1 pixel ≈ 5.87 millimeters of ground distance;

[0102] Based on excavator samples provided from the drone's perspective, a multimodal large model is used to generate excavator examples. If sample images with corresponding perspectives already exist, the excavator portion needs to be cropped from the sample images to ensure the inspection target occupies more than 50% of the entire image. Stable Diffusion Models typically generate images based on semantic information in a low-dimensional feature space. Therefore, existing image generation methods need to ensure the generated target is large enough, resulting in poor generation performance for small targets. For excavator samples, a Stable Diffusion Model is used to process color and texture variations. Through a conditional guidance mechanism, multiple appearance variations of the target object are generated, enriching the diversity of samples and enhancing the robustness of the model's recognition performance.

[0103] The Stable Diffusion Model was used to process the excavator samples for color and texture variations, generating multiple appearance variations of the inspection target, such as... Figures 2-4 As shown;

[0104] The generated excavator samples are quickly segmented using the SAM large model to rapidly obtain image masks, thus completely separating the excavator samples from the original background. Figure 5 As shown;

[0105] Several background images covering different scenes, such as construction sites and field work areas, need to be prepared. The excavator target is scaled according to the GSD accuracy conversion and then pasted onto the background images to generate several multimodal synthetic samples. Typically, an excavator is 4–6 meters long and 1.5–2.5 meters wide. Based on the GSD accuracy conversion, in the drone-captured images, L… ratio ∈(0.06,0.15), W ratio ∈(0.03,0.1), the generated excavator is randomly scaled based on its aspect ratio and composited into several background images, such as... Figure 6 As shown;

[0106] The generated samples were divided into training, validation and test sets, and trained using the YOLOv11 model. The initial learning rate was set to 0.001 and the number of iterations was 400.

[0107] This invention also provides an electronic device including a central processing unit (CPU), which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) or loaded from a storage unit into a random access memory (RAM). The RAM may also store various programs and data required for device operation. The CPU, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.

[0108] Multiple components in the device are connected to the I / O interface, including: input units such as keyboards and mice; output units such as various types of displays and speakers; storage units such as disks and optical discs; and communication units such as network interface cards (NICs), modems, and wireless transceivers. The communication unit allows the device to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0109] The processing unit performs the various methods and processes described above, such as the methods of the present invention. For example, in some embodiments, the methods of the present invention may be implemented as computer software programs tangibly contained in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program may be loaded and / or installed on the device via ROM and / or a communication unit. When the computer program is loaded into RAM and executed by the CPU, one or more steps of the methods of the present invention described above may be performed. Alternatively, in other embodiments, the CPU may be configured to execute the methods of the present invention by any other suitable means (e.g., by means of firmware).

[0110] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0111] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0112] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0113] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for generating multimodal image samples based on UAV perspective, characterized in that, The method includes: Identify the target objects for drone inspections and obtain information about those objects; Set the flight parameters for the UAV to perform the inspection mission, calculate the ground sampling distance GSD based on the flight parameters, and determine the actual ground size corresponding to each pixel in the image captured by the UAV; For cases where existing sample images exist and cases where no sample images exist, sample images of the target object are generated using a multimodal large model, respectively. The sample image of the generated target object is subjected to image segmentation processing to generate a segmentation mask, which completely separates the target object from the original background; Based on the information of the GSD and the target object, the separated target object is scaled proportionally and synthesized with multiple background images to generate multimodal image samples that conform to the real inspection scenario.

2. The method for generating multimodal image samples based on UAV perspective according to claim 1, characterized in that, The flight parameters for the UAV to perform inspection tasks include flight altitude, shooting angle, flight speed, and sensor information carried by the UAV.

3. The method for generating multimodal image samples based on UAV perspective according to claim 2, characterized in that, The sensor information carried by the drone includes image pixel L p ×W p a) Pixel size information, lens focal length, and lens type.

4. The method for generating multimodal image samples based on UAV perspective according to claim 1, characterized in that, The formula for calculating the ground sampling distance GSD is: Where f is the focal length of the camera lens, a is the pixel size, and H is the flight altitude.

5. The method for generating multimodal image samples based on UAV perspective according to claim 1, characterized in that, The specific components of generating sample images of the target object include: When an existing sample image exists, the sample image is cropped to ensure that the target object being inspected occupies more than 50% of the entire image. When no existing sample images exist, the target objects for drone inspection are generated using a multimodal large model based on the samples provided from the drone's perspective. Using a multimodal large model based on Diffusion Model, and through a conditional guidance mechanism, the target object after truncation and the target object generated by the multimodal large model are processed for color and texture changes, generating multiple appearance variations of the target object.

6. The method for generating multimodal image samples based on UAV perspective according to claim 1, characterized in that, The image segmentation process uses the SAM image segmentation model to automatically identify target objects in the image and generate accurate segmentation masks, completely separating the target objects from the original background and outputting independent target object materials for subsequent background compositing.

7. The method for generating multimodal image samples based on UAV perspective according to claim 4, characterized in that, The scaling specifically refers to: Based on the actual size and GSD value of the target object, calculate the proportion of the target object in the entire image, and scale the target object accordingly: Where L p W represents the number of horizontal pixels in the drone's image. p L represents the number of vertical pixels in the drone's image. obj W represents the actual length of the target object. obj L represents the actual width of the target object. ratio W represents the length of the target object after scaling. ratio This indicates the width of the target object after scaling.

8. The method for generating multimodal image samples based on UAV perspective according to claim 1, characterized in that, The multiple background images include: Images collected from actual inspection scenarios, public datasets, or computer-generated virtual scenes cover a variety of weather conditions, different terrains, and different light intensities.

9. The method for generating multimodal image samples based on UAV perspective according to claim 1, characterized in that, The method also includes the following steps: The generated image samples are divided into training set, validation set and test set; Training and validation were performed using the YOLO object detection model.

10. A method for generating multimodal image samples based on a UAV perspective according to claim 9, characterized in that, After the model is validated, the trained model will be deployed to the UAV inspection system to perform real-time target object detection.

11. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1 to 10.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 10.

Citation Information

Patent Citations

  • Method and device for generating defect image sample, medium and program product

    CN112967248A

  • Power grid unmanned aerial vehicle inspection image defect intelligent identification self-learning training method and system

    CN112990335A

  • Power transmission line channel abnormal crack sample generation method and system

    CN113240767A

  • Insulator defect detection method and system

    CN114693665A

  • Land area measurement method based on image segmentation network, unmanned aerial vehicle and medium

    CN116295134A