Method and system for generating aerial photography data of fine-grained multi-view unmanned aerial vehicle based on controllable graphics model
Patent Information
- Application Number
- CN202510266592.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-11
Smart Images

Figure CN120298518A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular, to a fine-grained multi-view UAV aerial photography data generation method and system based on a controllable text-to-image model. Background Art
[0002] With the development of the "low-altitude economy", unmanned aerial vehicles (UAVs) are increasingly being used in various smart city applications, such as logistics distribution, traffic detection, and police situation warning. In these tasks, UAVs need to accurately understand the urban scene. Due to its significantly better generalization performance than traditional vision algorithms, computer vision technology based on deep learning has gradually become one of the key supporting technologies for UAV environmental perception. However, when using computer vision algorithms based on deep learning to detect and segment targets in aerial photography scenes, there are two difficulties: one is that due to factors such as variable shooting angles, heights, distances, and complex low-altitude environments, there are differences in perspectives, postures, and appearances among similar targets; and there are significant differences in appearances among different types of targets. Therefore, it poses a great challenge to the robustness of vision algorithms; the other is that the sample data from the UAV perspective is insufficient, resulting in the visual model being prone to overfitting. Due to factors such as no-fly regulations and short battery life, the cost of collecting UAV data is relatively high. Compared with the natural image dataset taken from the planar perspective, the dataset scale from the UAV perspective is smaller. In addition, due to fixed shooting perspectives and heights, all real situations cannot be traversed, resulting in insufficient data diversity, and there are problems such as limited target categories, limited scenes, cluttered perspectives, missing perspectives, and heterogeneous data sources. The targets to be detected in the dataset show characteristics of uneven scales, uneven spatial distributions, uneven sample quantities, and uneven class semantics. Summary of the Invention
[0003] To solve some or all of the above technical problems existing in the prior art, the present invention provides a fine-grained multi-view UAV aerial photography data generation method and system based on a controllable text-to-image model.
[0004] The technical solution of the present invention is as follows:
[0005] In the first aspect, a fine-grained multi-view UAV aerial photography data generation method based on a controllable text-to-image model is provided. The method includes:
[0006] Create foreground target models with different postures through a virtual modeling tool, and collect image samples covering multiple perspectives and rotation angles to generate prior knowledge of the postures of the foreground targets;
[0007] Construct three-dimensional city models of multiple scene categories, and collect background scene data through different trajectories, heights, and perspectives;
[0008] Fuse the foreground target and the background scene to generate a layout structure diagram including customized aerial photography difficult example samples, where the aerial photography difficult example samples include crowded people, complex occlusion, and small target scenes;
[0009] Utilize a pre-trained image generation model, combined with the guiding conditions of the layout structure and weather conditions, to generate high-quality drone aerial photography image samples.
[0010] In an embodiment of the present invention, the generation of the pose prior knowledge of the foreground target specifically includes:
[0011] Use the Blender virtual engine to create three-dimensional digital target models, where the targets include pedestrians, two-wheeled vehicles, and four-wheeled vehicles;
[0012] Set up fine-grained multi-view virtual cameras to collect image samples of each target model at different perspectives and rotation angles, and generate the pose prior knowledge of the foreground target.
[0013] In an embodiment of the present invention, the construction of three-dimensional city models of multiple scene categories and the collection of background scene data through different trajectories, heights, and perspectives specifically include:
[0014] Use the Blender virtual engine to construct multiple urban scenes, including commercial blocks, residential blocks, and urban suburbs;
[0015] Design different shooting trajectories, select different shooting points on the shooting trajectories, and collect background scene data at different heights and perspectives at each shooting point along the shooting trajectories.
[0016] In an embodiment of the present invention, the fusion of the foreground target and the background scene to generate a layout structure diagram including aerial photography difficult example samples specifically includes:
[0017] Use a pre-trained segmentation model to segment the reasonable regions in the background scene map;
[0018] According to the target layout rules, superimpose the foreground target into the reasonable regions in the background scene to generate fusion samples, and the layout rules include crowded people scene generation rules, complex occlusion scene generation rules, and small target scene generation rules.
[0019] In an embodiment of the present invention, the generation of aerial photography image samples based on a pre-trained image generation model, combined with the guiding conditions of the layout structure and weather conditions specifically includes:
[0020] Utilize a controllable generation network to generate a semantic segmentation map to control the layout structure of the generated image;
[0021] Control the light, season, and weather conditions of the generated image through text guiding words;
[0022] Generate high-quality aerial image samples by combining the layout structure and text constraints of weather conditions.
[0023] In one embodiment of the present invention, the use of a controllable generation network to generate a semantic segmentation map to control the layout structure of the generated image specifically includes:
[0024] Automatically segment the fused image using a large semantic segmentation model to generate a semantic segmentation map;
[0025] Extract layout features at different scales and superimpose these features on the encoder of the pre-trained text-to-image large model to complete layout editing.
[0026] In one embodiment of the present invention, the control of the light, season, and weather conditions of the generated image by text guidance words specifically includes:
[0027] Design text guidance words including season, light, and weather conditions;
[0028] Use the pre-trained CLIP model to embed the text constraint conditions into the image space to generate weather condition constraint conditions.
[0029] In one embodiment of the present invention, the combination of the layout structure and text constraints of weather conditions to output high-quality images that meet the layout and weather constraints specifically includes:
[0030] Use the pre-trained variational autoencoder VAE and Unet network to construct a stable diffusion image generation model;
[0031] Perform denoising operations on the input image in the low-dimensional latent space, and combine the layout structure and text constraints of weather conditions to generate high-quality aerial image samples.
[0032] In a second aspect, a fine-grained multi-view unmanned aerial vehicle aerial photography data generation system based on a controllable text-to-image model is provided. The system includes:
[0033] A foreground target generation module for creating foreground target models in different poses through virtual modeling tools, collecting image samples covering multiple views and rotation angles, and generating prior knowledge of the poses of the foreground targets;
[0034] A background scene generation module for constructing three-dimensional city models of multiple scene categories and collecting background scene data through different trajectories, heights, and views;
[0035] A foreground-background fusion module that fuses the foreground targets and the background scenes to generate a layout structure diagram including customized aerial photography difficult example samples, and the aerial photography difficult example samples include crowded people, complex occlusions, and small target scenes;
[0036] An aerial photography difficult example sample generation module for generating high-quality drone aerial photography image samples by using a pre-trained image generation model and combining the guiding conditions of layout structure and weather conditions.
[0037] In an embodiment of the present invention, the aerial photography difficult example sample generation module includes:
[0038] A layout editing sub-module for generating a semantic segmentation map by using a controllable generation network to control the layout structure of the generated image;
[0039] A weather condition editing sub-module for controlling the light, season, and weather conditions of the generated image through text guiding words;
[0040] An image generation sub-module for generating high-quality aerial photography image samples by combining the text constraints of layout structure and weather conditions.
[0041] The main advantages of the technical solution of the present invention are as follows:
[0042] The fine-grained multi-view drone aerial photography data generation method and system based on the controllable text-to-image model of the present invention generate foreground targets with multiple views and poses and diverse background scene data through a virtual modeling tool, and generate high-quality aerial photography image samples by combining the constraints of layout structure and weather conditions, significantly improving the diversity and fine-grained characteristics of the data, supplementing difficult example scene samples, solving the problem that traditional methods are difficult to cover complex aerial photography conditions, not only reducing the data acquisition cost, but also enhancing the generalization ability and robustness of the visual perception model in complex scenes. Description of the Drawings
[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0044] Figure 1 It is a flowchart of the fine-grained multi-view drone aerial photography data generation method based on the controllable text-to-image model in an embodiment of the present invention;
[0045] Figure 2 It is the overall architecture of the fine-grained multi-view drone aerial photography data generation system based on the controllable text-to-image model in an embodiment of the present invention;
[0046] Figure 3 It is a schematic diagram of collecting foreground target images in the fine-grained multi-view drone aerial photography data generation method and system based on the controllable text-to-image model in an embodiment of the present invention;
[0047] Figure 4 Schematic diagram of the foreground-background fusion process in the method and system for generating fine-grained multi-view UAV aerial photography data based on a controllable text-to-image model according to an embodiment of the present invention;
[0048] Figure 5 Architecture diagram of the difficult example sample generation module in the method and system for generating fine-grained multi-view UAV aerial photography data based on a controllable text-to-image model according to an embodiment of the present invention;
[0049] Figure 6 Architecture diagram of the layout editing sub-module based on a controllable generation network in the method and system for generating fine-grained multi-view UAV aerial photography data based on a controllable text-to-image model according to an embodiment of the present invention;
[0050] Figure 7 Architecture diagram of the image generation sub-module based on graph-text constraints in the method and system for generating fine-grained multi-view UAV aerial photography data based on a controllable text-to-image model according to an embodiment of the present invention. Detailed implementation manners
[0051] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with specific embodiments of the present invention and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0052] The following will describe in detail the technical solutions provided by the embodiments of the present invention with reference to the drawings.
[0053] In a first aspect, an embodiment of the present invention provides a method for generating fine-grained multi-view UAV aerial photography data based on a controllable text-to-image model. As shown in the appendix Figure 1 The method includes:
[0054] S1, creating foreground target models with different poses through a virtual modeling tool, collecting image samples covering multiple views and rotation angles, and generating pose prior knowledge of the foreground targets.
[0055] Create foreground target models in the UAV aerial photography task through a virtual modeling tool. The target types include common pedestrians, vehicles, etc. For each target, set a virtual camera to perform circumferential shooting at different shooting angles and heights, so as to obtain image samples covering multiple views and rotation angles. These samples are refined into pose prior knowledge of the foreground targets for flexible invocation during subsequent data generation.
[0056] This step reduces the need for actual aerial photography sampling through virtual modeling, thereby lowering the acquisition cost. By creating foreground target models in different poses and then collecting image samples from multiple perspectives and rotation angles, the diversity of foreground target samples is increased, enhancing the realism and applicability of the generated data.
[0057] S2, Construct three-dimensional city models for various scene categories and collect background scene data at different trajectories, heights, and perspectives.
[0058] This step reduces the need for actual aerial photography sampling by virtually modeling three-dimensional city models, thereby lowering the acquisition cost. By presetting virtual shooting paths at different heights, angles, and trajectories, diverse background image data is collected, addressing the limitation that traditional data augmentation methods cannot comprehensively cover all background scenes.
[0059] S3, Fuse the foreground targets and background scenes to generate a layout structure diagram including customized aerial photography difficult example samples, where the aerial photography difficult example samples include scenes of crowded people, complex occlusions, and small targets.
[0060] Combining the collected foreground target and background scene data, a layout structure diagram is generated by defining reasonable layout rules, including difficult example scenes such as crowded people, complex occlusions, and small targets. These scenes are difficult to obtain in actual aerial photography but are extremely important when training a drone vision perception model.
[0061] Through this step, specific support for key application scenarios is provided, such as monitoring in high-density areas and small target detection, increasing the complexity and diversity of the dataset, enabling the training model to better handle difficult example scenes.
[0062] S4, Utilize a pre-trained image generation model and combine the guiding conditions of the layout structure and weather conditions to generate high-quality drone aerial photography image samples.
[0063] Utilize a pre-trained large image generation model (such as Stable Diffusion) to generate high-quality aerial photography images under the constraints of layout and weather text conditions. The layout structure provides spatial constraints, while the weather condition text controls the seasonal, lighting, and weather characteristics of the image. Designed in this way, it supports the generation of samples covering more weather and lighting conditions, significantly improving the quality and realism of the generated images and adapting to diverse application requirements.
[0064] The fine-grained multi-view UAV aerial photography data generation method based on a controllable text-to-image model provided by the embodiments of the present invention generates foreground targets and diverse background scene data with multiple views and poses through a virtual modeling tool, and combines the constraints of layout structure and weather conditions to generate high-quality aerial photography image samples, significantly improving the diversity and fine-grained characteristics of the data, supplementing difficult example scene samples, solving the problem that traditional methods are difficult to cover complex aerial photography conditions, not only reducing the data acquisition cost, but also enhancing the generalization ability and robustness of the visual perception model in complex scenes.
[0065] In some alternative embodiments of the present invention, in step S1 above, generating the pose prior knowledge of the foreground target specifically includes:
[0066] S11, Use the Blender virtual engine to create three-dimensional digital target models, including pedestrians, two-wheeled vehicles, and four-wheeled vehicles. For pedestrian targets, construct common human postures, such as walking, standing, running, squatting, lying down, etc. For vehicle targets: Design vehicle models of different types (such as cars, trucks, bicycles, motorcycles, etc.), and include state changes (such as door opening / closing, passenger-carrying / empty).
[0067] S12, Set up a fine-grained multi-view virtual camera to collect image samples of each target model at different views and rotation angles, and generate the pose prior knowledge of the foreground target. Through the fine-grained multi-view virtual camera, the target is photographed in a circular motion at different views (0° - 90°) and rotation angles (0° - 360°) to obtain image samples covering multiple views, different shooting heights, and directions. Then, refine the shooting data into the pose prior knowledge of the foreground target, including the appearance characteristics, proportional characteristics, and light and shadow changes of the target at different angles.
[0068] This embodiment provides target samples covering multiple views and pose changes, enhancing the diversity and fine-grained characteristics of the data set. And the method of modeling and collecting images significantly reduces the difficulty and cost compared with actual UAV sampling, especially when special pose targets (such as fallen pedestrians, vehicle accidents) are required, and the efficiency advantage is obvious.
[0069] In some alternative embodiments of the present invention, in step S2 above, constructing three-dimensional city models of multiple scene categories and collecting background scene data through different trajectories, heights, and views specifically includes:
[0070] S21, Use the Blender virtual engine to construct multiple urban scenes, including commercial blocks, residential blocks, and suburban areas of the city. Among them, the commercial block simulates high-density buildings, crowded roads, and various vehicle and pedestrian activities. The residential block contains low-rise buildings, wide sidewalks, and greening scenes. The suburban area of the city covers open spaces, sparse buildings, and natural environments.
[0071] S22. Design different shooting trajectories, select different shooting points on the shooting trajectories, and collect background scene data at different heights and angles of view along the shooting trajectories. Among them, when setting the virtual shooting trajectory, it covers the multi-level aerial photography requirements from low altitude to high altitude. For example, it can range from 5 meters to 1000 meters with an increment of 10 meters. At each shooting point on the trajectory, adjust the shooting direction (such as an inclination angle of 0° - 90°) and the rotation angle (a 0° - 360° rotation) to collect background data.
[0072] The background scene image acquisition method provided in this embodiment avoids problems such as no-fly restrictions, weather impacts, and high costs existing in actual acquisition, and can simulate complex background scenes, covering more urban environments and shooting conditions, enhancing the wide applicability of training data.
[0073] In some optional embodiments of the present invention, in the above step S3, fuse the foreground target and the background scene to generate a layout structure diagram including aerial photography difficult example samples, which specifically includes:
[0074] S31. Use a pre-trained segmentation model to segment the reasonable regions in the background scene map. The reasonable regions are the regions where the foreground targets can be placed (such as roads, sidewalks, open spaces, etc.).
[0075] S32. According to the target layout rules, superimpose the foreground target into the reasonable regions in the background scene to generate fusion samples. The layout rules include the generation rules for crowded scenes, complex occlusion scenes, and small target scenes. The layout rules clearly limit and define crowded scenes, complex occlusions, and small target scenes from multiple dimensions such as target density, occlusion situation, and target size. Exemplarily, for crowded scenes: simulate the distribution of a large number of people (number ≥ 10) in a specific scene (such as a commercial street). For complex occlusions: include mutual occlusion between targets (such as pedestrians and vehicles), occlusion between targets and the background (such as buildings and trees occluding targets). The occlusion degree is divided into mild (<20%), moderate (20% - 40%), and severe (40% - 60%). For small target scenes: generate samples where the target occupies less than 10% of the image, simulating the small target situation in long-distance aerial photography. The above is only an example of a layout rule provided in the embodiments of the present invention. Those skilled in the art can adaptively set the layout rules according to the application scenarios of the generated data set to meet the usage requirements, and the embodiments of the present invention do not make specific limitations in this regard.
[0076] Through the above operations, challenging scenes (such as high-density crowds and complex occlusions) are provided, enriching the aerial photography difficult example samples of the training data, and providing support for the performance improvement of the visual perception model.
[0077] In some alternative embodiments of the present invention, step S4, based on a pre-trained image generation model, generates aerial image samples in combination with guiding conditions of the layout structure and weather conditions, specifically including:
[0078] S41, Use a controllable generation network to generate a semantic segmentation map to control the layout structure of the generated image. By inputting layout structure information (such as a semantic segmentation map) into the controllable generation network, embed the layout features into the image generation model to ensure that the generated image conforms to the expected layout. Among them, the layout features include the position, size, density of foreground objects, and the reasonable distribution of the background scene.
[0079] S42, Control the light, season, and weather conditions of the generated image through text guiding words. Design text description conditions to specify the weather, light, and season features of the generated image. For example, "Aerial image of a city on a sunny spring day" or "Aerial image of a city street with fog at night". Use the pre-trained CLIP model to embed the text conditions and generate weather guiding conditions that match the description.
[0080] S43, Combine the text constraints of the layout structure and weather conditions to generate high-quality aerial image samples. Combine the dual constraints of the layout structure and weather conditions, and use the Stable Diffusion XL model to generate high-quality aerial image samples. The image generation process is achieved by gradually denoising, and integrates layout and weather information to ensure the diversity and quality of the samples.
[0081] Generating high-quality aerial images through this process provides precise control over the image layout and weather conditions, covers a variety of complex environments and light conditions, makes the generated image samples closer to the actual scene, and enhances the practicality of the dataset.
[0082] In some alternative embodiments of the present invention, the above-mentioned controlling the layout structure of the generated image through a semantic segmentation map specifically includes:
[0083] Use a large semantic segmentation model to automatically segment the fused image to generate a semantic segmentation map. Use a pre-trained semantic segmentation model (such as the Segment Anything Model, SAM) to segment the initial image generated by fusing the foreground object and the background scene, and extract the semantic segmentation map in the image. The segmentation map marks different region types, such as foreground objects (pedestrians, vehicles) and background objects (buildings, roads, etc.).
[0084] Extract layout features at different scales and superimpose these features onto the encoder of the pre-trained text-to-image large model to complete layout editing. Use methods such as Pixel-unshuffle to downsample the semantic segmentation map and convert the segmented image into feature maps of different scales. The extracted feature maps are embedded into the encoder of the image generation model in a layer-by-layer superimposed manner to ensure that the layout information is retained during the generation process.
[0085] This embodiment ensures the rationality of the layout of foreground objects and background scenes in the generated image, avoiding situations where the positions of foreground objects are illogical. And it improves the authenticity and consistency of the generated image. Especially in complex scene layouts, it can meet the requirements of actual drone tasks.
[0086] In some alternative embodiments of the present invention, controlling the light, season, and weather conditions of the generated image through the text guidance words specifically includes:
[0087] Design text guidance words that include seasons, light, and weather conditions. Design text descriptions according to weather changes, seasonal characteristics, and light conditions. For example: "Aerial images of a city on a sunny day in autumn" or "Aerial images of a city on a rainy night in winter". The text descriptions contain various combinations: 4 seasons (spring, summer, autumn, winter), 2 types of light (daytime, night), and 4 types of weather (sunny, rainy, snowy, foggy).
[0088] Use the pre-trained CLIP model to embed the text constraint conditions into the image space to generate weather condition constraint conditions. Use the CLIP model to embed the text description into the image generation model and associate the image and text modalities through the cross-attention mechanism. The embedded text features guide the generation process to make the generated image conform to the specified weather and light conditions.
[0089] This embodiment supports the generation of aerial image samples in various specific environments by flexibly adjusting the weather and light conditions. The generated data samples are more suitable for the environmental complexity in actual drone tasks and enhance the generalization ability of the training model.
[0090] In some alternative embodiments of the present invention, combining the text constraints of the layout structure and weather conditions to output high-quality images that meet the layout and weather constraints specifically includes:
[0091] Use the pre-trained variational autoencoder (VAE) and Unet network to construct a stable diffusion image generation model. Adopt the pre-trained stable diffusion image generation model (Stable Diffusion XL), whose structure includes a variational autoencoder (VAE) and a Unet network. The VAE module compresses the input image into a low-dimensional latent space and processes the image in this latent space to improve the computational efficiency.
[0092] Denoise the input image in the low-dimensional latent space, and combine the text constraints of the layout structure and weather conditions to generate high-quality aerial image samples. Starting from the initial random noise, gradually remove the noise by combining the layout structure features (generated from the semantic segmentation map) and the text conditions of the weather conditions. In each step of the denoising process, optimize according to the layout structure and weather conditions, and finally output high-quality aerial image samples.
[0093] In a second aspect, an embodiment of the present invention provides a fine-grained multi-view unmanned aerial vehicle (UAV) aerial photography data generation system based on a controllable text-to-image model, as shown in the appendix Figure 2 The system includes:
[0094] A foreground target generation module, configured to create foreground target models in different poses through a virtual modeling tool, collect image samples covering multiple views and rotation angles, and generate prior knowledge of the poses of the foreground targets;
[0095] A background scene generation module, configured to construct three-dimensional city models of multiple scene categories, and collect background scene data through different trajectories, heights, and views;
[0096] A foreground-background fusion module, which fuses the foreground targets and the background scenes to generate a layout structure diagram including customized difficult aerial photography example samples, and the difficult aerial photography example samples include crowded people, complex occlusions, and small target scenes;
[0097] A difficult aerial photography example sample generation module, configured to use a pre-trained image generation model, and combine the guiding conditions of the layout structure and weather conditions to generate high-quality UAV aerial photography image samples.
[0098] The fine-grained multi-view UAV aerial photography data generation system provided by the embodiment of the present invention generates foreground targets with multiple views and multiple poses and diverse background scene data through a virtual modeling tool, and combines the constraints of the layout structure and weather conditions to generate high-quality aerial photography image samples, significantly improving the diversity and fine-grained characteristics of the data, supplementing difficult example scene samples, and solving the problem that traditional methods are difficult to cover complex aerial photography conditions. It not only reduces the data acquisition cost, but also enhances the generalization ability and robustness of the visual perception model in complex scenes.
[0099] In some optional embodiments of the present invention, as shown in the appendix Figure 2 The difficult aerial photography example sample generation module includes:
[0100] A layout editing sub-module, configured to use a controllable generation network to generate a semantic segmentation map to control the layout structure of the generated image;
[0101] A weather condition editing sub-module, configured to control the light, season, and weather conditions of the generated image through text guiding words;
[0102] An image generation sub-module for generating high-quality aerial image samples by combining the text constraints of the layout structure and weather conditions.
[0103] The following is a detailed description of the fine-grained multi-view UAV aerial photography data generation method and system provided by the present invention in combination with specific embodiments.
[0104] Due to the implementation of the no-fly zone order, it is costly to use UAVs to collect aerial photography data, and it is difficult to collect sample data covering various urban scenes, complex interference situations, and weather conditions for training visual perception models. In the urban low-altitude scene, the targets under the UAV perspective include two categories: foreground targets and background targets. Foreground targets are pedestrians, two-wheeled vehicles, four-wheeled vehicles, etc.; background targets include static environmental targets such as roads, trees, and buildings. The present invention constructs a multi-view fine-grained aerial photography target dataset based on a controllable text-to-image model, generates foreground targets and background scenes respectively for combination, and obtains supplementary and enhanced data containing different urban scenes, interference situations, and weather conditions for model training of downstream target detection, semantic segmentation, and other environmental perception tasks.
[0105] The overall architecture of the fine-grained multi-view UAV aerial photography data generation method and system based on a controllable text-to-image model is as Figure 2 shown, mainly including a foreground target generation module, a background scene generation module, a foreground-background fusion module, and an aerial photography hard example sample generation module.
[0106] (1) Foreground target generation module
[0107] Compared with the planar perspective, it is difficult to describe and model the prior knowledge of the target appearance under the UAV perspective because the intra-class / inter-class targets have large morphological differences when the shooting perspective, height, and their own postures are different. In order to generate more representative aerial photography sample data and improve the diversity of the generated data, a foreground target generation module is designed. This module is mainly used to create various aerial photography target pose maps with different perspectives and postures for subsequent fusion with the generated background map to generate various aerial photography hard example sample data, such as sample data in bad weather conditions, dense crowd sample data, small target sample data, etc. These sample data will be used to train the on-board target detection model to improve the recognition and generalization performance of the model for various aerial photography targets under different weather conditions and shooting parameters.
[0108] Taking the generation of foreground human targets as an example, the foreground target generation module first uses the Blender virtual engine to create 3D digital human models in n common postures such as walking, standing, running, squatting, lying on the back, and crawling as pose metadata; on this basis, the joint angles in each pose metadata are adjusted to obtain m different pose variant digital human models for each meta-pose, and the total number of pose digital human models is. After completing the construction of the pose digital human models, fine-grained multi-view virtual cameras are set in the virtual engine environment to collect image samples of each digital human model from different perspectives and rotation angles, which are summarized as pose prior knowledge. Specifically, in each pose, the shooting distance is fixed, such as Figure 3 As shown, a visual knowledge acquisition hemisphere is created, and the shooting perspective and shooting height are adjusted for surround shooting to obtain the visual information of all perspectives of the target. The shooting perspective starts from 0 degrees of the frontal view and gradually increases by 30 degrees to 90 degrees of the top-down view; at each shooting perspective, starting from 0 degrees, it gradually increases by 30 degrees to 360 degrees to perform surround shooting on the target, and a set of image samples collected by surround shooting at each shooting perspective. Finally, a total of 19 shooting perspectives are collected for each pose of the digital human model, and each perspective contains 72 prior visual knowledge image samples of surround perspectives. Similarly, for bicycle, motorcycle, and car targets, k different models of each category are created. For the car model, two poses of door opening and door closing are designed; for the bicycle and motorcycle models, two poses of carrying passengers and not carrying passengers are designed. Similar to the visual knowledge acquisition method of the digital human model, for each vehicle pose, image samples from different perspectives and rotation angles are collected to extract visual prior knowledge.
[0109] (2) Background scene generation module
[0110] Due to the implementation of the no-fly order, it is difficult to ensure that during the process of collecting UAV aerial photography data, background image data in different scene categories, weather conditions, and light conditions can be comprehensively collected. In order to cover diverse background situations, a background scene generation module is designed. This module uses the Blender virtual engine to produce a three-dimensional city model, including i scenes such as commercial blocks, residential blocks, and urban suburbs. After completing the scene construction, different shooting trajectories are designed to collect scene background data along the trajectories at different heights and shooting perspectives. Specifically, starting from a shooting height of 5 meters, it gradually increases by 10 meters to 1000 meters for shooting. Shooting points are set at certain intervals on the shooting trajectory. Starting from 0 degrees at each shooting point, it gradually increases by 30 degrees to 90 degrees, and at each perspective, it performs 360-degree surround shooting starting from 0 degrees with an increment of 30 degrees to collect image data of different perspectives and different surround angles, and obtain prior visual knowledge image samples of the background scene.
[0111] (3) Foreground and background fusion module
[0112] This module is used to extract different foreground objects and background scenes for fusion. The purpose of fusion is to create some special aerial photography difficult example samples in the aerial photography dataset, such as: crowded people, small objects, complex occlusions, etc. If these aerial photography difficult example samples are collected by using a drone for aerial photography, due to the special situation, it is difficult to collect a sufficient number of samples. This module creates different aerial photography difficult example samples by setting corresponding rules for creating aerial photography difficult example samples, which are used as the image input of the subsequent aerial photography difficult example sample generation module, adding layout structure constraints to the image generation process.
[0113] The creation rules for crowded people samples include two variables: scene and number of people. The number of individuals in the crowded people should be greater than or equal to 10, and they appear in 3 scenes such as commercial areas, residential areas, and suburbs; the creation rules for complex occlusions include two variables: occlusion category and occlusion degree. The occlusion category is of two types: mutual occlusion between foreground-foreground objects and mutual occlusion between foreground-background objects. The occlusion degree is divided into three levels: mild occlusion, moderate occlusion, and severe occlusion. The area of mild occlusion is 0% - 20%, the area of moderate occlusion is 20% - 40%, and the area of severe occlusion is 40% - 60%. The creation rules for small object samples include two variables: scene and distance. The target pixel area accounts for 10% or less of the whole image, and it appears in 3 scenes such as commercial areas, residential areas, and suburbs. As Figure 4 shown, the fusion method is to first select an appropriate background scene image and foreground objects under the same shooting perspective, use a pre-trained segmentation model to segment reasonable areas such as roads and sidewalks in the background scene image where foreground objects can be added, and finally overlay the foreground objects onto the background scene according to the sample fusion rules to obtain a fused sample as the layout structure diagram of the aerial photography difficult example sample, which is used as the layout constraint for the subsequent generation process.
[0114] (4) Aerial photography difficult example sample generation module
[0115] Based on the previous module, under the constraints of layout structure and weather conditions, this method constructs an aerial photography difficult example sample generation module based on the pre-trained text-to-image large model Stable Diffusion XL to generate aerial photography image samples with different difficult example situations, seasonal situations, light conditions, and weather conditions. These samples will be used to supplement and expand the existing aerial photography dataset, construct a fine-grained drone aerial photography visual knowledge dataset, and be used to train downstream visual perception models such as object detection and semantic segmentation. As Figure 5 shown, the aerial photography difficult example sample generation module mainly consists of: a layout editing sub-module based on a controllable generation network, a weather condition editing sub-module based on text embedding, and an image generation sub-module based on graph-text constraints.
[0116] 1) Layout editing sub-module based on a controllable generation network
[0117] This module creates layout constraints through a controllable generation network to control the layout structure of the final generated image. The structure of the controllable generation network is as shown in Figure 6 . Its function is to align the prior knowledge inside the pre-trained text-to-image large model with the external generation control instructions, so as to more precisely control the final generation result of the model. The input of this module is the semantic segmentation map S0 of the scene, and this segmentation map will be used as a constraint condition to control the composition layout of the final generated image:
[0118] S0 = SAM(X0)
[0119] where X0 is the output image of the foreground-background fusion module, and SAM(·) is the semantic segmentation large model SegmentAnything Model. After automatically segmenting the fused image using the SAM model, the Pixel-unshuffle method is used to downsample the segmented image to a resolution of 64×64, and then 4 feature extraction modules are constructed to extract layout features of different scales These conditional features will be added to the feature maps of each level of the Unet network encoder in the pre-trained text-to-image large model in a superimposed manner as follows:
[0120]
[0121] where C Layout is the output feature map with layout constraints, which is used to embed layout information in the subsequent image generation process to complete the layout editing of the final generated image.
[0122] 2) Weather condition editing sub-module based on text embedding
[0123] In addition to layout constraints, a text description embedding module is further constructed to control the lighting conditions, season conditions, and weather conditions of the final generated image in a text-guided manner. The advantage of text constraints is that they can improve the diversity of generated images through very precise descriptions. This module first designs text guiding words (Prompt) as the constraint conditions for the subsequent image generation sub-module, and then uses the pre-trained CLIP model to complete the embedding (Embedding) of the text constraint conditions in the image space. The text guiding words include 4 seasons of spring, summer, autumn, or winter, 2 lighting conditions of day or night, and 4 weather conditions of sunny, rainy, snowy, or foggy, such as: "Generate an aerial image of a city by a drone in the morning on a sunny day in winter". There are a total of 32 combinations of weather condition conditions, which can be used as text constraint conditions for the subsequent image generation module.
[0124] Let the text input be t, and the cross-attention mechanism is used to establish the association between the text and image modalities to complete the embedding of the text constraint conditions, as follows:
[0125]
[0126] Among them, U(·) is the Unet encoder, T(·) is the pre-trained CLIP model, and W Q , W K , W V is a mapping matrix.
[0127] 3) Image generation sub-module based on graph-text constraint
[0128] The basic structure of the image generation sub-module based on graph-text constraint is the Stable Diffusion XL (SDXL) model, which is pre-trained using the large dataset LAION-5B. LAION-5B contains 5 billion text-image samples. By fitting this dataset, the SDXL model has accumulated a large amount of prior knowledge for text-to-image generation. The present invention improves based on this model and designs a stable diffusion image generation model based on aerial photography prior knowledge (UAV-image Stable Diffusion XL model, UAV-SDXL model). By embedding the layout structure picture and the text description of the weather condition through the previous module, graph-text constraint conditions are formed to guide the image generation process of the UAV-SDXL model, and finally, a high-quality and diverse dataset of difficult cases for UAV aerial photography is generated using the prior knowledge inside the model.
[0129] As Figure 7 shown, the UAV-SDXL model consists of two parts: a pre-trained Variational Autoencoder (VAE) and Unet. The VAE module contains an encoder and a decoder. The encoder part downsamples the input image X to Z0 in the low-dimensional latent space, and then the decoder restores the image. The reason for using the VAE module for downsampling is that operating on the image in the low-dimensional latent space can effectively improve the operation efficiency. The Unet part of the model mainly performs denoising operations on the input image in the low-dimensional latent space, and the denoising optimization formula is as follows:
[0130]
[0131] Among them, represents the noise map at the t-th step, C Layout is the layout structure constraint condition, T Condition is the weather condition constraint condition, ε θIt represents the noise predicted by Unet at the t-th step. Using this formula, noise is first added to Z0 to obtain the true noise value at each step. During training, the weight values of the pre-trained VAE and Unet parts in the UAV-SDXL model are frozen, and only the weights of each layer of the controllable generation network are optimized. Then, during the image generation process, the initial noise feature map Z is sampled from an arbitrary Gaussian distribution T , and based on this, denoising optimization is carried out. In each step of denoising, with C Layout and T Condition as the constraint conditions, Unet is used to generate the predicted noise ε θ , and then Z T is subtracted by ε θ , and enter the next step of denoising. After T rounds of denoising process, a low-dimensional feature map without noise is obtained and sent to the decoder of the VAE module to output the final generated result.
[0132] The inference process after training is as follows: Input the text guidance word of the weather condition into the text embedding module, input the initial image fused by the foreground and background into the controllable generation network, use the UAV-SDXL model to gradually denoise the randomly sampled noise, and finally generate a drone aerial image with a specified layout and weather condition.
[0133] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. In addition, "front", "back", "left", "right", "up", and "down" in this article are all referenced based on the placement state shown in the drawings.
[0134] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A fine-grained multi-view UAV aerial photography data generation method based on a controllable text-to-image model, characterized in that, Including: Create foreground target models in different postures through virtual modeling tools, collect image samples covering multiple perspectives and rotation angles, and generate prior knowledge of the postures of foreground targets; Construct 3D city models of multiple scene categories, and collect background scene data through different trajectories, heights and perspectives; Fuse the foreground target and the background scene to generate a layout structure diagram including customized aerial photography difficult example samples, and the aerial photography difficult example samples include crowded people, complex occlusion and small target scenes; Use a pre-trained image generation model, combined with the guiding conditions of the layout structure and weather conditions, to generate high-quality UAV aerial photography image samples.
2. The method for generating fine-grained multi-view UAV aerial photography data based on a controllable text-to-image model according to claim 1, wherein The generation of prior knowledge of the posture of the foreground target specifically includes: Use the Blender virtual engine to create 3D digital target models, and the targets include pedestrians, two-wheeled vehicles and four-wheeled vehicles; Set up fine-grained multi-perspective virtual cameras to collect image samples of each target model at different perspectives and rotation angles, and generate prior knowledge of the postures of foreground targets.
3. The method for generating fine-grained multi-view UAV aerial photography data based on a controllable text-to-image model according to claim 1, characterized in that The construction of 3D city models of multiple scene categories and the collection of background scene data through different trajectories, heights and perspectives specifically include: Use the Blender virtual engine to construct multiple urban scenes, including commercial blocks, residential blocks, and urban suburbs; Design different shooting trajectories, select different shooting points on the shooting trajectories, and collect background scene data at different heights and perspectives along the shooting trajectories at each shooting point.
4. The method for generating fine-grained multi-view UAV aerial photography data based on a controllable text-to-image model according to claim 1, wherein, The fusion of the foreground target and the background scene to generate a layout structure diagram including aerial photography difficult example samples specifically includes: Use a pre-trained segmentation model to segment the reasonable regions in the background scene map; According to the target layout rules, superimpose the foreground target into the reasonable regions in the background scene to generate fusion samples, and the layout rules include crowded people scene generation rules, complex occlusion scene generation rules, and small target scene generation rules.
5. The method for generating fine-grained multi-view UAV aerial photography data based on a controllable text-to-image model according to claim 1, wherein Based on the pre-trained image generation model, combined with the guiding conditions of the layout structure and weather conditions to generate aerial photography image samples specifically includes: Use a controllable generation network to generate a semantic segmentation map to control the layout structure of the generated image; Control the light, season and weather conditions of the generated image through text guiding words; Combine the text constraints of the layout structure and weather conditions to generate high-quality aerial photography image samples.
6. The method for generating fine-grained multi-view UAV aerial photography data based on a controllable text-to-image model according to claim 5, wherein, The use of a controllable generation network to generate a semantic segmentation map to control the layout structure of the generated image specifically includes: Use a large semantic segmentation model to automatically segment the fused image to generate a semantic segmentation map; Extract layout features of different scales, and superimpose these features into the encoder of the pre-trained text-to-image large model to complete layout editing.
7. The method for generating fine-grained multi-view UAV aerial photography data based on a controllable text-to-image model according to claim 5, wherein The control of the light, season and weather conditions of the generated image through text guiding words specifically includes: Design text guiding words including season, light, and weather conditions; Use the pre-trained CLIP model to embed the text constraint conditions into the image space to generate weather condition constraint conditions.
8. The method for generating fine-grained multi-view UAV aerial photography data based on a controllable text-to-image model according to claim 5, wherein The combination of the text constraints of the layout structure and weather conditions to output high-quality images that meet the layout and weather constraints specifically includes: Use the pre-trained variational autoencoder VAE and Unet network to construct a stable diffusion image generation model; Denoise the input image in the low-dimensional latent space, and generate high-quality aerial image samples by combining the text constraints of the layout structure and weather conditions.
9. A fine-grained multi-view UAV aerial photography data generation system based on a controllable text-to-image model, characterized in that, Including: A foreground object generation module, which is used to create foreground object models with different postures through virtual modeling tools, collect image samples covering multiple perspectives and rotation angles, and generate prior knowledge of the postures of foreground objects; A background scene generation module, which is used to construct 3D city models of multiple scene categories and collect background scene data through different trajectories, heights and perspectives; A foreground and background fusion module, which fuses the foreground object and the background scene to generate a layout structure diagram including customized aerial photography difficult case samples, and the aerial photography difficult case samples include crowded people, complex occlusion and small target scenes; An aerial photography difficult case sample generation module, which is used to generate high-quality UAV aerial photography image samples by using a pre-trained image generation model and combining the guiding conditions of the layout structure and weather conditions.
10. The fine-grained multi-view UAV aerial photography data generation system based on a controllable text-to-image model according to claim 9, characterized in that, The aerial photography difficult case sample generation module includes: A layout editing sub-module, which is used to generate a semantic segmentation map by using a controllable generation network to control the layout structure of the generated image; A weather condition editing sub-module, which is used to control the light, season and weather conditions of the generated image through text guiding words; An image generation sub-module, which is used to generate high-quality aerial photography image samples by combining the text constraints of the layout structure and weather conditions.
Citation Information
Cited By
Multi-modal image sample generation method and device based on view angle of unmanned aerial vehicle, and medium
CN120894645A