Aerial image data set construction method based on simulation rendering technology
By combining ODM and Blender to generate aerial image datasets containing real geographic information and virtual features, the problems of high cost, low efficiency and lack of diversity in aerial dataset production are solved. This achieves efficient and diverse dataset generation and automated annotation, which is suitable for a variety of computer vision and remote sensing tasks.
Patent Information
- Application Number
- CN202511625250.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-07
- Publication Date
- 2026-02-03
AI Technical Summary
Existing technologies suffer from high production costs, low efficiency, limited diversity, difficulty in annotation, and a disconnect between real and virtual data, making it difficult to meet the demand for high-quality datasets.
By combining the open-source aerial image processing tool ODM and the 3D modeling and rendering software Blender, an aerial image dataset containing real geographic information and controllable virtual elements is generated through simulation rendering technology. An automated annotation method is used to integrate real terrain data and virtual scenes to generate diverse aerial image datasets.
It significantly reduces costs, improves efficiency, enhances dataset diversity and annotation accuracy, and generates datasets containing real geographic coordinates and pose information, suitable for a variety of computer vision and remote sensing tasks, and supports deep learning model training.
Smart Images

Figure CN121458901A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of computer vision, UAV remote sensing image processing and 3D reconstruction technology, and specifically relates to a method for constructing an aerial image dataset based on simulation rendering technology. Background Technology
[0002] With the popularization of drone technology and the advancement of computer vision algorithms, high-quality aerial image datasets have become crucial in precision agriculture, smart cities, disaster emergency response, environmental monitoring, and the training and evaluation of artificial intelligence, especially deep learning models. Currently, acquiring such datasets primarily relies on in-situ drone flight data collection; however, this method has significant limitations: High cost and long cycle: Organizing flights, obtaining permits, and processing data are time-consuming, labor-intensive, and costly. Because the target is difficult to transport, missions are sometimes impossible to complete.
[0003] Environmental and conditional limitations: Due to factors such as weather, airspace control, seasonal changes, and specific events such as disaster accessibility, it is difficult to obtain data under diverse scenarios and conditions, such as different lighting, weather, and seasons.
[0004] Labeling difficulties: Fine-grained labeling of real aerial images, such as object detection boxes and semantic segmentation masks, relies heavily on manual work, which is inefficient, costly, and prone to errors.
[0005] Insufficient data diversity: Field data collection cannot exhaust all possible scenario changes, target types, and anomalies.
[0006] ODM, as an open-source aerial image processing toolchain, can reconstruct 3D scenes based on overlapping real aerial images through motion reconstruction and multi-view stereo technology, outputting geospatial data products including sparse / dense point clouds, digital land models, orthophoto maps, and 3D mesh models. ODM is widely used in geographic information systems and surveying, but its core function is limited to processing real images; it cannot actively generate new aerial images, nor can it simulate unrealistic or specific conditions such as heavy rain, dense fog, or nighttime shooting effects.
[0007] Blender, a powerful open-source 3D creation suite, boasts comprehensive capabilities in modeling, materials, lighting, animation, and rendering. It can accurately simulate drone flight paths and camera parameters, generating realistic aerial images and easily adding various virtual objects with automated annotation. However, Blender-generated scenes are typically virtual constructs, lacking realistic geographic coordinates and precise terrain features, resulting in a gap between the generated data and real-world geospatial data. Summary of the Invention
[0008] To address the problems of high production cost, low efficiency, limited diversity, difficult annotation, and separation of real and virtual data in existing aerial photography datasets, this invention provides a method for constructing aerial image datasets based on simulation rendering technology. It integrates the advantages of the open-source aerial image processing tool ODM (OpenDroneMap) in real geographic information processing with the controllability of the open-source 3D modeling and rendering software Blender in virtual scene generation and precise annotation. This method can efficiently construct aerial image datasets that simultaneously contain real geographic information and controllable virtual elements, and supports structured annotation, improving the diversity and annotation accuracy of the dataset. It is suitable for various application scenarios such as artificial intelligence model training, remote sensing analysis, and 3D visualization.
[0009] The objective of this invention is achieved through the following technical solution: This invention provides a method for constructing an aerial image dataset based on simulation rendering technology, comprising the following steps: Step S1: In the aerial image processing tool ODM, generate terrain data based on real aerial images; Step S2: Import the terrain data generated in step S1 into Blender software to construct a virtual aerial scene that integrates the real terrain; Step S3: In the virtual aerial photography scene constructed in step S2, drive the simulated drone to perform simulated drone aerial photography at preset sampling points, and generate virtual aerial photography images containing metadata and virtual aerial photography images containing target object annotations respectively. Step S4: Use ODM to verify the virtual aerial image containing metadata generated in step S3. If the verification is successful, proceed to step S5; otherwise, return to step S1. Step S5: Organize the raw data, process data, virtual aerial images containing metadata, and virtual aerial images containing target object annotations to obtain an aerial image dataset.
[0010] Further, step S1 includes: S11: Input a sequence of real aerial images with overlap, captured by a drone, into the aerial image processing tool ODM; S12: Use ODM to process the real aerial image sequence to obtain terrain data; S13: Export the terrain data generated by the ODM.
[0011] Further, step S2 includes: S21: Import the terrain data exported in step S1 into the Blender software environment to generate a terrain model; S22: Based on the terrain model generated in step S21, add the required virtual object models to obtain the initial virtual aerial photography scene; S23: Set materials and textures for objects in the initial virtual aerial scene; S24: Configure the lighting system for the initial virtual aerial photography scene to simulate different times and weather conditions; S25: Construct a target object labeling system for the types of virtual objects that need to be identified in the initial virtual aerial photography scene, and obtain a virtual aerial photography scene that integrates the real terrain.
[0012] Further, step S3 includes: S31: Simulate drone parameter settings and camera parameter settings for a single preset sampling point; S32: Generate virtual aerial images of a drone from multiple perspectives at preset sampling points, and simultaneously execute steps S33 and S34 respectively; S33: Synchronously record metadata and generate virtual aerial images containing metadata; S34: Automatically annotate target objects in multi-view virtual aerial images to generate virtual aerial images containing target object annotations; S35: Select the next preset sampling point and repeat steps S31 to S34 to obtain virtual aerial images containing metadata and virtual aerial images containing target object annotations for all sampling points.
[0013] Further, step S31 includes: In the virtual aerial photography scene built in Blender, set the flight path, flight altitude, and flight speed of the simulated drone; Configure virtual camera parameters to simulate a real drone camera.
[0014] Further, step S33 includes: Drive the simulated drone to fly along a set path and trigger camera rendering at preset sampling points to generate virtual aerial images of the simulated drone from multiple perspectives; While rendering each virtual aerial image, the precise metadata corresponding to that image is recorded and stored, and the transportation bureau is exported as a structured file. Metadata is embedded into the EXIF information of the virtual aerial image, and then the virtual aerial image is saved as a standard image format to obtain a virtual aerial image containing metadata.
[0015] Further, step S34 includes: Perform initial settings for the camera and target object list; Define the camera movement range and set the save path for virtual aerial images and target object label files; Traverse all target objects, obtain the bounding boxes of the target objects within the virtual aerial image, and complete the automatic annotation of the target object information; Save the virtual aerial image with automatic target object annotation as a standard image format to generate a virtual aerial image containing target object annotations.
[0016] Further, step S5 includes: S51: Integrate the automatic annotation results from step S3; S52: The image data, metadata, automatic annotation files, real terrain data, virtual scene configuration files, and software processing logs from steps S1 to S3 are stored in a unified manner; the image data includes the real aerial images, virtual aerial images containing metadata, and virtual aerial images containing target object annotations. S53: Output aerial image dataset.
[0017] The present invention has the following advantages: Deep coupling of realism and virtuality: Using ODM to process real images to obtain accurate terrain as a basis, and combining it with Blender to add virtual elements and environmental changes, a hybrid dataset is generated that has both a real geographical basis and high controllability and diversity.
[0018] The data is highly diverse and controllable: it can flexibly adjust lighting, weather, season, flight parameters, camera parameters, as well as the type, quantity, and location of virtual objects to generate massive amounts of data covering various scenarios and conditions, overcoming the limitations of on-site collection.
[0019] High degree of automation and low cost in annotation: Blender's built-in labeling system can generate annotation information directly during rendering, significantly reducing manual annotation workload and improving annotation efficiency and accuracy. When annotating datasets, using Blender scripts for automatic annotation saves annotation time, especially when dealing with large amounts of data.
[0020] Data consistency and quality are guaranteed: The synthesized images are reverse-processed and verified through ODM to ensure that the generated virtual images conform to the optical and geometric imaging principles of the real world, thus guaranteeing the geometric consistency of the dataset used to train the model and forming a closed-loop quality control.
[0021] Supports multimodal tasks: The generated datasets not only contain images, but also include precise geographic locations, poses, camera parameters, and various forms of annotations, which can be directly used for a variety of computer vision and remote sensing tasks such as object detection, semantic segmentation, instance segmentation, depth estimation, 3D reconstruction, and UAV visual positioning and navigation.
[0022] Significantly reduces costs and increases efficiency: It reduces reliance on expensive and time-consuming field drone flights and can quickly generate large amounts of high-quality training data indoors, making it particularly suitable for training deep learning models that require large-scale data.
[0023] Complete geospatial information: The generated images come with real geographic coordinates and attitude information, making them directly applicable to geographic information systems and remote sensing analysis tasks that require geographic reference.
[0024] High scalability: Built on open source tools, it is easy to integrate new plugins, model libraries and annotation tools, and adapt to future changes in needs. Attached Figure Description
[0026] Figure 1 This is a flowchart of a method for constructing an aerial image dataset based on simulation rendering technology, as described in an embodiment of the present invention. Figure 2 This is a schematic diagram illustrating the ODM processing of real aerial images to generate terrain data in an embodiment of the present invention; Figure 3 This is a schematic diagram of the Blender software interface in an embodiment of the present invention; Figure 4 This is a schematic diagram of data collective annotation in an embodiment of the present invention. Detailed Implementation
[0028] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments: This embodiment describes a method for constructing an aerial image dataset based on simulation rendering technology, such as... Figure 1 As shown, it includes the following steps: Step S1: In the aerial image processing tool ODM, generate terrain data based on real aerial images: S11: Input a sequence of real aerial images with overlap, captured by a drone, into the aerial image processing tool ODM.
[0029] S12: Use ODM to process the real aerial image sequence to obtain terrain data; the image processing flow includes image feature matching, sparse point cloud reconstruction, dense point cloud generation, digital surface model generation, orthophoto map generation, and three-dimensional mesh model reconstruction.
[0030] S13: Export terrain data generated by ODM, including but not limited to: dense point cloud, digital surface model (DSM), orthophoto DOM, and 3D mesh model.
[0031] Step S2: Import the terrain data generated in Step S1 into Blender software to construct a virtual aerial scene that integrates the real terrain. S21: Import the terrain data exported in step S1 into the Blender software environment to generate a terrain model.
[0032] Preferably, after importing terrain data, the terrain model may be located far from the origin due to geographic coordinate offset. Therefore, necessary coordinate system transformation, scale adjustment, and position alignment should be performed during terrain data import. In Blender, enter object mode, select the model object, click the model to highlight it, and right-click to set the origin. Hover the mouse over the view window, press the right mouse button, and select Set Origin → Origin to Geometric Center.
[0033] S22: Based on the terrain model generated in step S21, add the required virtual object models, such as buildings, vegetation, vehicles, pedestrians, road facilities, etc., to obtain the initial virtual aerial photography scene.
[0034] S23: Set realistic materials and textures for objects in the initial virtual aerial scene.
[0035] S24: Configure the lighting system for the initial virtual aerial scene to simulate different times and weather conditions.
[0036] S25: Construct a target object labeling system for the types of virtual objects that need to be identified in the initial virtual aerial scene, and obtain a virtual aerial scene that integrates the real terrain: Assign unique identifiers or specific material channels to virtual object types that need to be identified in virtual aerial photography scenes, such as "cars," "pedestrians," "houses," "trees," and "roads." The automatic labeling system can obtain the object's position and dimensions through the unique identifiers or specific material channels.
[0037] Step S3: In the virtual aerial photography scene constructed in step S2, drive the simulated drone to perform simulated drone aerial photography at preset sampling points, generating virtual aerial images containing metadata and virtual aerial images containing target object annotations respectively: S31: Simulate drone parameter settings and camera parameter settings for a single preset sampling point: In the virtual aerial photography scene built by Blender, set the flight path (track point sequence or flight route), flight altitude, and flight speed of the simulated drone; Configure virtual camera parameters to simulate a real drone camera, including: focal length, sensor size, resolution, optical distortion coefficient, and exposure parameters.
[0038] S32: Generate virtual aerial images of a drone from multiple perspectives at preset sampling points, and then execute steps S33 and S34 respectively; S33: Synchronously record metadata and generate virtual aerial images containing metadata: Drive the simulated drone to fly along a set path and trigger camera rendering at preset sampling points to generate virtual aerial images of the simulated drone from multiple perspectives; While rendering each virtual aerial image, the precise metadata corresponding to that image is recorded and stored, and exported by the Transportation Bureau as a structured file; the metadata includes: camera position, camera intrinsic parameters, timestamp, key points (to ensure that the metadata is accurately associated with the virtual aerial image). Metadata is embedded into the EXIF information of the virtual aerial image, and then the virtual aerial image is saved as a standard image format to obtain a virtual aerial image containing metadata.
[0039] S34: Automatically annotate target objects in multi-view virtual aerial images to generate virtual aerial images containing target object annotations: Initialize the camera and target object list settings: First, initialize the scene, camera, rotation axis, lights and target object list, where the target object is the target object of the virtual object model added in step S25; Define the camera movement range and set the save path for virtual aerial images and target object label files; The system iterates through all target objects, obtaining their bounding boxes within the virtual aerial image, thus automatically annotating the target objects. Specifically, it uses a camera transformation matrix to project the vertices of the target objects onto screen space. If the object is not within the camera's field of view, it returns "None"; otherwise, it returns a string of annotation information, which consists of bounding box coordinates formatted in YOLO format. Furthermore, manual annotation of the target objects is also possible, or the automatic annotation results can be manually corrected.
[0040] Save the completed virtual aerial image, which has been automatically or manually annotated / corrected, as a PNG file to generate a virtual aerial image containing the target object annotations.
[0041] S35: Select the next preset sampling point and repeat steps S31 to S34 to obtain virtual aerial images containing metadata and virtual aerial images containing target object annotations for all sampling points.
[0042] Step S4: Verify the virtual aerial image containing metadata generated in Step S3 using the ODM. If the verification is successful, proceed to Step S5; otherwise, return to Step S1. S41: Input the virtual aerial image sequence containing metadata generated in step S3 into the ODM system.
[0043] S42: Process the virtual aerial image input in step S41 using the standard ODM workflow to obtain ODM reconstructed terrain data.
[0044] S43: Verification: Check whether the ODM processing was completed successfully; The geometric consistency of the ODM reconstructed terrain data is compared with the original real terrain data obtained in step S1 (i.e., the terrain data imported into Blender). If the geometric consistency comparison passes, proceed to step S5 to prepare for integrating the aerial image dataset.
[0045] Step S5: Organize the raw data (referring to the real aerial images used in Step S1 and the terrain data obtained by ODM), process data (referring to the terrain model generated in Step S2, the initial virtual aerial scene, and the virtual aerial scene fused with real terrain), virtual aerial images containing metadata, and virtual aerial images containing target object annotations to obtain an aerial image dataset: S51: Integrate the automatic annotation results from step S3 with manual annotation / correction information.
[0046] S52: Establish a dataset version management system to uniformly store, track versions, and manage the association of image data, metadata, automatically labeled files, real terrain data, virtual scene configuration files, and software processing logs. The image data includes the real aerial images, virtual aerial images containing metadata, and virtual aerial images containing target object annotations.
[0047] S53: Output the final structured aerial image dataset, which includes image data, automatically labeled files, metadata, real terrain data, terrain reference files (terrain reference files used to generate terrain data in step S1), and system documentation (referring to the software system documentation generated by the dataset version management system established in step S52).
Claims
1. A method for constructing an aerial image dataset based on simulation rendering technology, characterized in that, Includes the following steps: Step S1: In the aerial image processing tool ODM, generate terrain data based on real aerial images; Step S2: Import the terrain data generated in step S1 into Blender software to construct a virtual aerial scene that integrates the real terrain; Step S3: In the virtual aerial photography scene constructed in step S2, drive the simulated drone to perform simulated drone aerial photography at preset sampling points, and generate virtual aerial photography images containing metadata and virtual aerial photography images containing target object annotations respectively. Step S4: Use ODM to verify the virtual aerial image containing metadata generated in step S3. If the verification is successful, proceed to step S5; otherwise, return to step S1. Step S5: Organize the raw data, process data, virtual aerial images containing metadata, and virtual aerial images containing target object annotations to obtain an aerial image dataset.
2. The method for constructing an aerial image dataset based on simulation rendering technology as described in claim 1, characterized in that, Step S1 includes: S11: Input a sequence of real aerial images with overlap, captured by a drone, into the aerial image processing tool ODM; S12: Use ODM to process the real aerial image sequence to obtain terrain data; S13: Export the terrain data generated by the ODM.
3. The method for constructing an aerial image dataset based on simulation rendering technology as described in claim 1, characterized in that, Step S2 includes: S21: Import the terrain data exported in step S1 into the Blender software environment to generate a terrain model; S22: Based on the terrain model generated in step S21, add the required virtual object models to obtain the initial virtual aerial photography scene; S23: Set materials and textures for objects in the initial virtual aerial scene; S24: Configure the lighting system for the initial virtual aerial photography scene to simulate different times and weather conditions; S25: Construct a target object labeling system for the types of virtual objects that need to be identified in the initial virtual aerial photography scene, and obtain a virtual aerial photography scene that integrates the real terrain.
4. The method for constructing an aerial image dataset based on simulation rendering technology as described in claim 1, characterized in that, Step S3 includes: S31: Simulate drone parameter settings and camera parameter settings for a single preset sampling point; S32: Generate virtual aerial images of a drone from multiple perspectives at preset sampling points, and simultaneously execute steps S33 and S34 respectively; S33: Synchronously record metadata and generate virtual aerial images containing metadata; S34: Automatically annotate target objects in multi-view virtual aerial images to generate virtual aerial images containing target object annotations; S35: Select the next preset sampling point and repeat steps S31 to S34 to obtain virtual aerial images containing metadata and virtual aerial images containing target object annotations for all sampling points.
5. The method for constructing an aerial image dataset based on simulation rendering technology as described in claim 4, characterized in that, Step S31 includes: In the virtual aerial photography scene built in Blender, set the flight path, flight altitude, and flight speed of the simulated drone; Configure virtual camera parameters to simulate a real drone camera.
6. The method for constructing an aerial image dataset based on simulation rendering technology as described in claim 4, characterized in that, Step S33 includes: Drive the simulated drone to fly along a set path and trigger camera rendering at preset sampling points to generate virtual aerial images of the simulated drone from multiple perspectives; While rendering each virtual aerial image, the precise metadata corresponding to that image is recorded and stored, and the transportation bureau is exported as a structured file. Metadata is embedded into the EXIF information of the virtual aerial image, and then the virtual aerial image is saved as a standard image format to obtain a virtual aerial image containing metadata.
7. The method for constructing an aerial image dataset based on simulation rendering technology as described in claim 4, characterized in that, Step S34 includes: Perform initial settings for the camera and target object list; Define the camera movement range and set the save path for virtual aerial images and target object label files; Traverse all target objects, obtain the bounding boxes of the target objects within the virtual aerial image, and complete the automatic annotation of the target object information; Save the virtual aerial image with automatic target object annotation as a standard image format to generate a virtual aerial image containing target object annotations.
8. The method for constructing an aerial image dataset based on simulation rendering technology as described in claim 1, characterized in that, Step S54 includes: S51: Integrate the automatic annotation results from step S3; S52: The image data, metadata, automatic annotation files, real terrain data, virtual scene configuration files, and software processing logs from steps S1 to S3 are stored in a unified manner; the image data includes the real aerial images, virtual aerial images containing metadata, and virtual aerial images containing target object annotations. S53: Output aerial image dataset.