UAV-Based Realistic 3D Automatic Modeling Method and System
Through the drone real-life three-dimensional automation modeling method, using technologies such as motion fuzzy convolution suppression, panoramic geometric splicing and deep semantic recognition, the problems of low automation and low accuracy in the existing technology are solved, efficient and automated three-dimensional modeling is achieved, and the realism and accuracy of the model is improved, and it is suitable for virtual reality and augmented reality applications.
Patent Information
- Application Number
- CN202510294611.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-03-13
AI Technical Summary
The existing three-dimensional modeling technology of drone relies on manual intervention and manual post-processing, which has problems such as low data processing efficiency, low accuracy and insufficient automation. Especially in complex environments, it is difficult to process massive data and identify different scene elements, and it is difficult to provide accurate three-dimensional data in dynamic scenes or difficult-to-get detailed areas.
Using drone-based real-life three-dimensional automated modeling method, aerial video data is automatically processed through motion fuzzy convolution suppression processing, panoramic geometric stitching, pixel-by-pixel-deep semantic recognition, dynamic light rendering evolution, vegetation type fine-grained recognition and fuzzy area detection and other technologies, combined with deep learning and physical light rendering, aerial video data is automatically processed to build a high-precision three-dimensional panoramic model.
It realizes efficient and automated three-dimensional modeling, improves image clarity and modeling accuracy, enhances the realism and immersion of the model, reduces manual intervention, and is suitable for applications such as virtual reality and augmented reality.
Smart Images

Figure CN119810359B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of 3D modeling, and particularly to a method and system for automatic 3D modeling of real scenes based on unmanned aerial vehicles (UAVs). Background Art
[0002] With the continuous development of technology and the increasing demand for industrial applications, 3D modeling technology has been widely used in various fields, especially in urban planning, architectural design, archaeological site reconstruction, natural landscape analysis, etc. Traditional 3D modeling methods usually rely on ground surveys, manual photography, and complex computer-aided design software. These methods are not only cumbersome to operate, but also have poor environmental adaptability. Especially in the modeling of large-scale complex scenes, there are high time costs and personnel inputs.
[0003] In recent years, with the rapid development of UAV technology, UAV aerial photography, as a new data acquisition method, has been widely used in the 3D modeling of various actual scenes. By equipping UAVs with high-definition camera devices or light detection and ranging (LiDAR) sensors, it is possible to quickly obtain large-scale and high-precision images and point cloud data, greatly improving the efficiency and accuracy of modeling. However, traditional UAV-based 3D modeling methods still face some challenges. Especially in complex environments, how to effectively process massive data, identify different scene elements, and perform accurate modeling remains a difficult point.
[0004] In addition, most existing UAV 3D modeling technologies rely on manual intervention and manual post-processing, resulting in problems such as low data processing efficiency, low accuracy, and insufficient automation. For some dynamic scenes or difficult-to-obtain detailed areas, traditional modeling methods often cannot provide accurate 3D data. To address these problems, there is an urgent need in modern industry and research fields for a more intelligent and efficient method for automatic 3D modeling of real scenes based on UAVs. Summary of the Invention
[0005] To solve the above technical problems, the present invention proposes a method and system for automatic 3D modeling of real scenes based on UAVs to solve at least one of the above technical problems.
[0006] To achieve the above object, the present invention provides a method for automatic 3D modeling of real scenes based on UAVs, including the following steps:
[0007] Step S1: Obtain the UAV aerial video; perform motion blur convolution suppression processing and panoramic geometric stitching on the UAV aerial video to construct a panoramic stitching reconstruction diagram;
[0008] Step S2: Perform pixel-by-pixel depth semantic recognition on the panoramic stitching reconstruction diagram to generate urban building areas and natural environment areas;
[0009] Step S3: Perform dynamic light rendering evolution and analysis of the physical form structure of urban buildings, and construct a 3D urban building rendering model;
[0010] Step S4: Conduct fine-grained identification of vegetation types in the natural environment area, and perform 3D panoramic fusion based on the 3D urban building rendering model to construct a 3D panoramic model;
[0011] Step S5: Detect blurred areas in the 3D panoramic model, and conduct low-altitude secondary aerial photography to extract high-definition videos of the blurred areas at low altitude;
[0012] Step S6: Optimize the local details of the blurred areas based on the high-definition videos of the blurred areas at low altitude, so as to construct a 3D real-scene optimization model.
[0013] Through motion blur convolution suppression processing, the system can real-time eliminate the negative impacts brought by blur, improve the clarity of the video, and enhance the accuracy of subsequent image analysis and modeling. The panoramic geometric stitching technology aligns and synthesizes multiple aerial video frames to generate a seamless and complete panoramic image. This not only avoids the problem of inconsistent multi-perspectives in the traditional shooting process, but also ensures that data from different shooting angles can complement each other, providing broader and more accurate scene information, and laying a foundation for subsequent semantic analysis and 3D modeling. Through per-pixel depth semantic recognition, the system can accurately distinguish and label the urban building and natural environment areas in the image. Using a deep learning model for pixel-level semantic segmentation can extract different elements such as buildings, roads, and vegetation in different scenarios, and then provide reliable semantic information for constructing a high-precision 3D model. The identified urban building areas and natural environment areas provide a clear boundary division for subsequent steps, enabling the modeling and processing of different areas to adopt the most suitable algorithms and technologies. The building areas can adopt fine-grained modeling, while the natural environment areas conduct terrain analysis and vegetation recognition. Through these precise area divisions, the modeling system can efficiently perform local optimization and integration in the subsequent stage. The changes in light and shadow have a profound impact on the visual effect of buildings. The dynamic light rendering evolution technology can simulate the light effects at different time periods and different weather conditions, generating realistic shadow projections and light reflections. Through this dynamic light change, the 3D urban model can be more vivid and real, providing high-quality visual performance. By analyzing the physical structure of buildings (such as wall thickness, roof shape, window position, etc.), the system extracts the geometric features of the building and conducts precise modeling. This process not only ensures the authenticity of the building appearance, but also can accurately reflect the actual structure and spatial layout of the building in the virtual environment. The vegetation in the natural environment area is usually diverse and widely distributed. Through the fine-grained vegetation recognition technology, the system can accurately identify different types of vegetation (such as trees, shrubs, lawns, etc.), and analyze its coverage area and distribution. This helps to construct a more realistic natural environment model and provides data support for subsequent applications such as landscape design and ecological assessment. By fusing the 3D models of urban building areas and natural environment areas, the system can create a comprehensive and accurate 3D panoramic scene. The seamless integration of urban buildings and the natural environment not only improves the visual effect of the model, but also enhances the sense of space and hierarchy of the scene, truly realizing the natural presentation of the virtual world. Even high-resolution aerial videos may still have blurred areas in some cases, affecting the accuracy of 3D modeling. Through the blur area detection technology, the system can automatically identify and label the blurred parts in the image. This technology not only improves the accuracy of modeling, but also reduces the need for later manual intervention. For the detected blurred areas, the system conducts secondary shooting by flying the drone at low altitude to obtain high-definition local detail images.This process can not only eliminate blurriness and enhance detail clarity, but also ensure that every detail in the 3D model is accurately restored, especially in the complex facades of buildings or the detailed parts of the natural environment. Through high-definition images captured by low-altitude videos, the system can optimize the details in the blurry areas. Using image enhancement techniques (such as super-resolution reconstruction, texture detail enhancement, etc.), the local details are finely processed, making the texture of the 3D model clearer and more realistic. The optimized model effectively improves the user experience, especially in applications such as virtual reality (VR) or augmented reality (AR), where detail optimization is crucial. By combining the locally optimized details from low-altitude videos, the constructed 3D real-scene optimized model will be more accurate and can present high-quality visual effects in various scenarios. This optimization makes the final 3D scene model not only geometrically accurate, but also more realistic in terms of lighting effects and material surfaces, and can perfectly restore the actual environment.
[0014] In this specification, a drone-based real-scene 3D automatic modeling system is provided for performing the drone-based real-scene 3D automatic modeling method as described above, including:
[0015] A geometric stitching module, configured to obtain drone aerial videos; perform motion blur convolution suppression processing and panoramic geometric stitching on the drone aerial videos, so as to construct a panoramic stitching reconstruction map;
[0016] A depth semantic recognition module, configured to perform per-pixel depth semantic recognition on the panoramic stitching reconstruction map, so as to generate urban building areas and natural environment areas;
[0017] A light rendering evolution module, configured to perform dynamic light rendering evolution and building physical form structure analysis on the urban building areas, and construct a 3D urban building rendering model;
[0018] A 3D panoramic fusion module, configured to perform fine-grained recognition of vegetation types in the natural environment areas, and perform 3D panoramic fusion according to the 3D urban building rendering model, so as to construct a 3D panoramic model;
[0019] A blurry area detection module, configured to detect blurry areas in the 3D panoramic model, and perform low-altitude secondary aerial photography to extract low-altitude high-definition videos of the blurry areas;
[0020] A local detail optimization module, configured to perform local detail optimization of the blurry areas according to the low-altitude high-definition videos of the blurry areas, so as to construct a 3D real-scene optimized model.
[0021] The present invention can effectively remove the blurring effect and restore image details by using motion blur convolution suppression algorithms (such as blind deblurring algorithms), ensuring that each frame of the image has a high-quality visual effect. This is particularly important for subsequent 3D modeling, as any degradation in image quality will affect feature extraction and model accuracy. Through panoramic geometric stitching technology, aerial images from multiple perspectives can be accurately stitched into a complete panoramic image. Geometric stitching minimizes the perspective distortion between different images and effectively synthesizes a seamless large-scale view, providing a high-precision image basis for subsequent depth semantic recognition and 3D reconstruction. Through deep learning networks (such as convolutional neural networks, Transformers, etc.), pixel-by-pixel depth semantic recognition can be performed to classify and label each element in the image at the pixel level. This not only helps to accurately identify elements such as buildings, roads, and vegetation, but also distinguishes the boundaries between natural environment areas and urban building areas based on environmental complexity, thereby improving the subsequent modeling accuracy. Physical lighting and dynamic rendering algorithms (such as ray tracing or physically based rendering PBR technology) are used to simulate the light changes at different times and angles. The impact of light on buildings, such as light reflection and shadow changes, is accurately reproduced, providing a sense of realism for the visual presentation of buildings under different lighting conditions. By simulating the evolution of light, the final building model appears more vivid and natural in the real environment. Analyzing the physical structure of the building (such as wall thickness, window frames, roof structure, etc.) can construct a more accurate 3D building model, avoiding the limitations of models based solely on flat drawings. This not only improves the modeling accuracy but also enhances the realism of the building, providing high-quality support for urban planning and virtual reality applications. Through deep learning algorithms (such as FCN, Mask R-CNN, etc.), fine-grained classification and recognition of vegetation types in the natural environment can be carried out. This process can be accurate to different types of plants, trees, etc., and can even identify the growth state of plants (such as healthy, withered, etc.), providing strong data support for the accurate modeling of the natural environment. Combining the modeling results of urban building areas and natural environment areas, the two are integrated into a whole through 3D panoramic fusion technology (such as seamless fusion of lighting, materials, and geometric models). In this way, the final constructed 3D panoramic model not only shows the architectural features of the city but also accurately reproduces the details of the surrounding natural environment, enhancing the realism and immersion of the entire scene. Through image processing technology, it is possible to identify blurred areas caused by distance, movement, etc. during aerial photography. This detection step is very important because blurred areas will affect the accuracy and integrity of modeling. Timely discovery and marking of these areas provide guidance for subsequent optimization. Once blurred areas are detected, low-altitude secondary aerial photography can be carried out for high-definition acquisition at a closer distance and a more stable flight state.This can obtain clearer and higher-resolution video data, making up for the loss of details in the first aerial photography and providing a high-quality data source for the subsequent repair of three-dimensional modeling details. Through the deep learning detail recognition of low-altitude videos, important information such as texture details and structural features in the blurred areas are extracted and applied to the optimization of the three-dimensional model. This not only improves the accuracy of the blurred areas but also increases the fineness and realism of the entire three-dimensional model. Based on the optimized detail information, a three-dimensional real-scene optimization model with high precision and quality is constructed. The three-dimensional model after detail repair has higher visual realism and spatial restoration degree and can be applied to multiple fields such as urban planning, virtual reality, and architectural design. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a schematic diagram of the step flow of a method for automatically building a three-dimensional real scene based on an unmanned aerial vehicle according to the present invention;
[0023] Figure 2 It is a schematic diagram of the detailed implementation step flow of step S1;
[0024] Figure 3 It is a schematic diagram of the detailed implementation step flow of step S2. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0026] The embodiments of the present application provide a method and system for automatically building a three-dimensional real scene based on an unmanned aerial vehicle. The execution subjects of the method and system for automatically building a three-dimensional real scene based on an unmanned aerial vehicle include, but are not limited to, mechanical equipment, data processing platforms, cloud server nodes, network upload devices, etc. that carry the system, which can be regarded as general computing nodes of the present application. The data processing platform includes, but is not limited to, at least one of an audio and image management system, an information management system, and a cloud data management system.
[0027] Please refer to Figures 1 to 3 , the present invention provides a method for automatically building a three-dimensional real scene based on an unmanned aerial vehicle. The method for automatically building a three-dimensional real scene based on an unmanned aerial vehicle includes the following steps:
[0028] Step S1: Obtain the aerial photography video of the unmanned aerial vehicle; perform motion blur convolution suppression processing and panoramic geometric stitching on the aerial photography video of the unmanned aerial vehicle to construct a panoramic stitching reconstruction diagram;
[0029] Step S2: Perform per-pixel depth semantic recognition on the panoramic stitching reconstruction diagram to generate urban building areas and natural environment areas;
[0030] Step S3: Perform dynamic light rendering evolution and building physical form structure analysis on urban building areas, and construct a 3D urban building rendering model;
[0031] Step S4: Conduct fine-grained identification of vegetation types in natural environment areas, and perform 3D panoramic fusion based on the 3D urban building rendering model to construct a 3D panoramic model;
[0032] Step S6: Detect blurred areas in the 3D panoramic model, and conduct low-altitude secondary aerial photography to extract low-altitude high-definition videos of the blurred areas;
[0033] Step S9: Optimize local details of the blurred areas based on the low-altitude high-definition videos of the blurred areas, thereby constructing a 3D real-scene optimization model.
[0034] It should be noted that there seems to be a numbering error in your original text. The step number in "Step S9" should probably be "S6" as per the previous sequence. I have translated it according to the corrected understanding. If this is not what you intended, please clarify.Through motion blur convolution suppression processing, the system can real-time eliminate the negative impact brought by blur, improve the clarity of the video, and enhance the accuracy of subsequent image analysis and modeling. The panoramic geometric stitching technology generates a seamless and complete panoramic image by aligning and synthesizing multiple aerial video frames. This not only avoids the problem of inconsistent multi-perspectives in the traditional shooting process, but also ensures that data from different shooting angles can complement each other, providing a broader and more accurate scene information, laying a foundation for subsequent semantic analysis and 3D modeling. Through per-pixel depth semantic recognition, the system can accurately distinguish and label the urban building and natural environment areas in the image. Using a deep learning model for pixel-level semantic segmentation can extract different elements such as buildings, roads, and vegetation in different scenes, and thus provide reliable semantic information for constructing a high-precision 3D model. The identified urban building areas and natural environment areas provide a clear boundary division for subsequent steps, enabling the modeling and processing of different areas to adopt the most suitable algorithms and technologies. Fine-grained modeling can be used for the building area, while terrain analysis and vegetation recognition are carried out for the natural environment area. Through these precise area divisions, the modeling system can efficiently perform local optimization and integration in the subsequent stage. The changes in light and shadow have a profound impact on the visual effect of buildings. The dynamic light rendering evolution technology can simulate the light effects at different time periods and different weather conditions, generating realistic shadow projections and light reflections. Through this dynamic light change, the 3D urban model can be more vivid and real, providing high-quality visual performance. By analyzing the physical structure of buildings (such as wall thickness, roof shape, window position, etc.), the system extracts the geometric features of the building and conducts accurate modeling. This process not only ensures the authenticity of the building appearance, but also can accurately reflect the actual structure and spatial layout of the building in the virtual environment. The vegetation in the natural environment area is usually diverse and widely distributed. Through fine-grained vegetation recognition technology, the system can accurately identify different types of vegetation (such as trees, shrubs, lawns, etc.), and analyze their coverage area and distribution. This helps to construct a more realistic natural environment model and provides data support for subsequent applications such as landscape design and ecological assessment. By fusing the 3D models of urban building areas and natural environment areas, the system can create a comprehensive and accurate 3D panoramic scene. The seamless integration of urban buildings and natural environment not only improves the visual effect of the model, but also enhances the sense of space and hierarchy of the scene, truly realizing the natural presentation of the virtual world. Even high-resolution aerial videos may still have blurred areas in some cases, affecting the accuracy of 3D modeling. Through the blur area detection technology, the system can automatically identify and label the blurred parts in the image. This technology not only improves the accuracy of modeling, but also reduces the need for later manual intervention. For the detected blurred areas, the system conducts secondary shooting by the low-altitude flight of the drone to obtain high-definition local detail images.This process can not only eliminate blurriness and enhance detail clarity, but also ensure that every detail in the 3D model is accurately restored, especially in the complex facades of buildings or the detailed parts of natural environments. Through the high-definition images captured by low-altitude videos, the system can optimize the details of the blurry areas. By using image enhancement techniques (such as super-resolution reconstruction, texture detail enhancement, etc.), the local details are finely processed, making the texture of the 3D model clearer and more realistic. The optimized model effectively improves the user experience. Especially in applications such as virtual reality (VR) or augmented reality (AR), detail optimization is crucial. By combining the locally optimized details from low-altitude videos, the constructed 3D real-scene optimized model will be more accurate and able to present high-quality visual effects in various scenarios. This optimization makes the final 3D scene model not only geometrically accurate, but also more realistic in terms of lighting effects and material surfaces, and can perfectly restore the actual environment.
[0035] In the embodiments of the present invention, refer to Figure 1 , which is a schematic diagram of the step flow of a method for automatic 3D real-scene modeling based on an unmanned aerial vehicle in the present invention. In this example, the steps of the method include:
[0036] Step S1: Obtain the aerial video captured by the unmanned aerial vehicle; perform motion blur convolution suppression processing and panoramic geometric stitching on the aerial video captured by the unmanned aerial vehicle, so as to construct a panoramic stitching reconstruction map;
[0037] In this embodiment, an unmanned aerial vehicle (UAV) is used to carry a high-definition imaging device to conduct aerial photography of a target area. The range of the aerial photography area should be determined according to subsequent modeling requirements. The flight route of the UAV should ensure that it can cover all angles of the required area to obtain sufficient image data. During the flight, the UAV is controlled to fly stably, ensuring that the camera lens always points to the target area and avoiding rapid movement as much as possible to reduce the generation of motion blur. An image restoration-based deconvolution method is adopted to restore the image by calculating the blur kernel in the video frame. The specific method is to use frequency domain filtering or spatial domain convolution inversion technology to estimate the blur kernel (such as the direction and degree of motion blur) and apply deconvolution processing to restore the details of the image. This method can effectively eliminate motion blur from the image and significantly improve the image clarity. For each video frame, first, frequency domain analysis of the blur is performed through an image processing algorithm (such as Fourier transform), and then the clear image is restored through a deconvolution algorithm. For motion blur, L2 regularization is used to optimize the deconvolution result and reduce noise and artifacts. A feature matching-based stitching algorithm (such as SIFT, SURF, or ORB algorithm) is used. By extracting key feature points in the image, matching and transformation correction are performed. First, feature points are extracted from each frame of the image, and then the RANSAC (Random Sample Consensus) algorithm is used to eliminate incorrect matching points and calculate the transformation matrix. These transformation matrices are applied to the image to achieve geometric transformation of the image. Finally, image stitching is performed through image fusion technology to ensure a natural transition between images and no obvious seams. There are often different exposure differences in the stitched images, which can lead to color differences at the seams. To solve this problem, multi-exposure fusion (MEF) or image blending technology (such as multi-resolution fusion technology) is used to smooth the brightness and color differences between images and generate a smooth and natural panoramic image.
[0038] Step S2: Perform pixel-by-pixel depth semantic recognition on the panoramic stitching reconstruction diagram to generate urban building areas and natural environment areas;
[0039] In this embodiment, U-Net or DeepLabv3+ is selected as the algorithm for deep semantic segmentation. These models perform excellently in pixel-level classification and are suitable for processing urban buildings and natural environment areas in complex scenarios. DeepLabv3+ is particularly good at capturing multi-scale features by introducing dilated convolution to increase the receptive field. The model is pre-trained using a high-quality semantic segmentation dataset, such as Cityscapes or ADE20K. These datasets contain rich annotations of urban and natural scenes and are suitable for the model to learn. Fine-tuning is performed on the selected dataset using the Adam optimizer, with a learning rate set to 0.0001, a batch size of 4, and the number of training epochs set to 100. During the training process, the loss and accuracy on the validation set are monitored to prevent overfitting. The panoramic stitching reconstruction image is preprocessed, including resizing the image (such as 512x512) to meet the input requirements of the model. The cv2.resize() function in OpenCV is used for image scaling. Normalization is performed to scale the pixel values to the range [0, 1] to improve the convergence speed of the model. The preprocessed panoramic image is input into the trained deep learning model to perform forward inference and generate class labels for each pixel. The model outputs a probability map with the same size as the input image, where each pixel corresponds to the probability of a different class. Threshold processing is performed on the probability map, with a threshold set to 0.5. Pixels with probabilities greater than this value are labeled as the corresponding class, such as buildings, roads, vegetation, etc. The recognition results are visualized, with each class using a different color. The cv2.applyColorMap() function in OpenCV is used to map the class labels to a visualized image. The processing results are saved to generate a labeled image and output as a file (such as PNG format) for subsequent analysis. The number of pixels of each class (urban building area and natural environment area) in the panoramic stitching image is counted to quantify the proportion of different areas. The np.unique() function in NumPy is used to achieve this. The analysis results are recorded to generate a report containing information such as the pixel proportion and area of each area.
[0040] Step S3: Perform dynamic light rendering evolution and building physical form structure analysis on the urban building area to construct a three-dimensional urban building rendering model;
[0041] In this embodiment, Unity or Unreal Engine is selected as the 3D rendering engine. These engines support high-quality dynamic lighting and real-time rendering, and are suitable for building complex urban building scenes. The Lumen lighting system of Unreal Engine can achieve global illumination and dynamic reflection, which is suitable for analyzing the ambient lighting changes of buildings. Multiple light sources are set in the scene, including directional lights (simulating sunlight), point lights, and spotlights. Adjust the intensity (such as Intensity = 3.0) and color of the light sources to simulate the lighting effects at different time periods. Set the attenuation range of the light sources to ensure that the lighting covers the entire building area and simulates the lighting changes in the real world. Through drone aerial video, collect the lighting data of the urban building area at different time periods, and calculate the reflection and shadow effects of each building under different lighting conditions. Use image processing technology to analyze the impact of lighting intensity and direction on the building surface, and generate time series data of lighting changes. Create a 3D model of the building in the rendering engine, and apply the materials and textures obtained from the lighting analysis. Use the dynamic lighting system to adjust the angle and intensity of the light sources in real time to simulate the evolution process of lighting. Record the rendering results at each time point to generate a lighting animation. Record the dynamic light rendering process as a video file, and compress it using an encoding format such as H.264 to ensure a balance between picture quality and file size. Output and visualize the lighting effects at different time periods to facilitate observing the performance of the building under lighting changes. Use Blender or Maya for the physical form analysis of the building. These tools support precise measurement and analysis of 3D models and are suitable for extracting the structural features of buildings. In the analysis tool, measure each part of the building, including wall thickness, opening size, top structure, etc. Set the measurement accuracy to 0.01 meters to ensure the accuracy of the data. Analyze the building layer by layer, and extract structural parameters such as wall thickness, window height, and number of floors from the bottom layer to the top layer. Record the data values of each parameter. Use an automated script (such as a Python script) to batch extract data in Blender to improve efficiency. Combine the dynamic light rendering results with the building physical form structure data to generate a complete 3D urban building rendering model. Use Unity or Unreal Engine to integrate all elements together. Perform final lighting and material adjustments in the rendering engine to ensure that the generated model is as close as possible to the real-world effect. Export the complete 3D urban building rendering model as a shareable 3D file format (such as FBX or OBJ) for use and display on other platforms.
[0042] Step S4: Conduct fine-grained identification of vegetation types in the natural environment area, and perform 3D panoramic fusion based on the 3D urban building rendering model to construct a 3D panoramic model;
[0043] In this embodiment, ResNet50 or EfficientNet is selected as the model for fine-grained vegetation type recognition. These models perform excellently in image classification tasks and are suitable for processing complex natural environment images. ResNet50 can effectively learn deep features through the introduction of residual connections and is used to identify different types of vegetation, such as trees, shrubs, lawns, etc. The model is pre-trained using publicly available datasets containing multiple vegetation types, such as PlantNet or Oxford Pets. These datasets contain a large number of labeled samples and are suitable for fine-grained classification. Fine-tuning is performed on the selected dataset using the Adam optimizer, with a learning rate set to 0.0001, a batch size of 16, and the number of training epochs set to 50. During the training process, the loss and accuracy on the validation set are monitored to ensure the generalization ability of the model. The images of natural environment areas are preprocessed, including resizing the images (such as 224x224) to meet the input requirements of the model. The cv2.resize() function in OpenCV is used for image scaling. Normalization processing is carried out to scale the pixel values to the range of [0, 1] to improve the convergence speed of the model. The preprocessed images are input into the trained model for forward inference to generate the vegetation type labels for each pixel. The model outputs a probability map with the same size as the input image, and each pixel corresponds to the probability of different vegetation types. Threshold processing is performed on the probability map, with the threshold set to 0.6, and the pixels with probabilities greater than this value are marked as the corresponding vegetation types. The recognition results are visualized, and different colors are used to mark each vegetation type. The cv2.applyColorMap() function in OpenCV is used to map the class labels to a visualized image. The processing results are saved, and the marked images are generated and output as files (such as PNG format) for subsequent analysis. Point cloud fusion and voxel fusion techniques are selected to fuse the vegetation information obtained from fine-grained recognition with the 3D urban building rendering model to generate a complete 3D panoramic model. Open3D or PCL (Point Cloud Library) is used for point cloud processing. These libraries support efficient point cloud data operations and visualization. The identified vegetation areas are integrated with the data of the 3D urban building rendering model to ensure that the coordinate systems of the two are consistent and avoid position misalignment. The 3D model of the vegetation area is generated, and relevant algorithms (such as Delaunay Triangulation) are used to convert the planar vegetation information into a 3D model. The create_from_point_cloud function in Open3D is used to convert the point cloud data of urban buildings and vegetation into a 3D model. The normal calculation parameters of the point cloud are set to improve the accuracy of model details. Downsampling and filtering are performed on the point cloud data, and the voxel grid method is used for noise reduction, with the voxel size set to 0.05 meters to ensure the smoothness of the point cloud.In a rendering engine (such as Unity or Unreal Engine), perform real-time rendering on the fused 3D model to ensure the realism of lighting and materials. Export the final 3D panoramic model into a shareable 3D file format (such as FBX or OBJ) for use and display on other platforms.
[0044] Step S5: Detect the blurred areas of the 3D panoramic model, conduct a second low-altitude aerial survey, and extract high-definition low-altitude videos of the blurred areas;
[0045] In this embodiment, a detection algorithm based on image sharpness is selected. Common detection metrics include Laplace transform and Sobel operator, and these two methods can effectively evaluate the clarity of an image. The Laplace transform calculates the second derivative of an image, and the transformation result of a blurred image is usually small. Therefore, a threshold is set to judge the degree of blurriness of the image. During the detection process, the blurriness threshold is set to 100 (the specific threshold can be adjusted according to the actual situation). If the result of the Laplace transform is lower than this value, the area is considered blurred. Load the image data generated by the 3D panoramic model, and use the cv2.imread() function of OpenCV to read the image. Perform grayscale processing on the image, and use the cv2.cvtColor() function to convert the image into a grayscale image. Apply the Laplace transform, and use the cv2.Laplacian() function to calculate the sharpness of the image. Compare the result with the set blurriness threshold to identify the blurred areas. Mark the detected blurred areas with a border or color, and use the cv2.rectangle() function to draw a rectangle on the original image to indicate the location of the blurred areas. Save the marked image for subsequent analysis and reference. Determine the aerial photography height and flight path, set the flight height to 30 meters to ensure that the details of the blurred areas can be clearly captured. Plan the aerial photography route to ensure that all blurred areas are covered. Adopt a grid flight method so that each blurred area can be fully photographed. Select a drone equipped with a high-resolution camera, such as DJI Mavic 3, which can shoot 4K (3840x2160 pixels) high-definition video to ensure the video quality. Set the camera parameters, select an appropriate shutter speed (such as 1 / 1000 second) and ISO value (such as 100) to adapt to different lighting conditions and ensure the image clarity. According to the planned aerial photography path, start the drone for low-altitude aerial photography, and monitor the flight status and shooting effect in real time to ensure that the drone flies at a safe height. Take multiple shots above each blurred area to ensure obtaining high-definition video data from multiple angles and perspectives. Save the high-definition video data generated during the aerial photography process to the memory card of the drone to ensure data backup to prevent loss. Ensure that the shooting duration of each blurred area is not less than 10 seconds for subsequent analysis and processing. Use video processing software (such as Adobe Premiere Pro or Final Cut Pro) to import the aerial photography video for editing. Clip out the video segments of the blurred areas captured, and ensure that the duration of each segment is the same for subsequent comparison and analysis. Export the processed low-altitude high-definition video, select an appropriate encoding format (such as H.264) to ensure that the video quality and file size are appropriate. Save it in a standard format (such as MP4) for subsequent display and analysis.
[0046] Step S6: Optimize the local details of the blurred areas based on the low-altitude high-definition videos of the blurred areas, so as to construct a 3D real-scene optimization model.
[0047] In this embodiment, a detail optimization method based on super-resolution reconstruction technology is selected, such as SRCNN (Super-Resolution Convolutional Neural Network) or GAN (Generative Adversarial Network). These methods can effectively enhance the details and clarity of blurred videos. SRCNN can recover more details in blurred areas by learning the mapping relationship between low-resolution images and high-resolution images. Use a publicly available dataset containing rich architectural and vegetation details (such as Set5 or DIV2K) for pre-training the model, set the learning rate to 0.0001, the batch size to 32, and the number of training epochs to 100. Set the super-resolution output size to 2 times that of the original video to increase the clarity of details. Use OpenCV or FFmpeg to load the low-altitude high-definition video of the blurred area to ensure frame-by-frame processing of the video frames. Extract the video frames into an image sequence, use the cv2.VideoCapture() function to read the video, and call cv2.imwrite() to save each frame as an image file. Preprocess each extracted frame image, including resizing the image to meet the input requirements of the super-resolution model (such as 33x33), and perform normalization processing to ensure that the pixel values are between [0, 1]. Input the preprocessed frame images into the trained super-resolution model, perform forward inference, and generate high-resolution output images. Record the super-resolution images of each frame and use the cv2.imwrite() function to save the processed images. Synthesize all the high-resolution images into a video, use the cv2.VideoWriter() function of FFmpeg or OpenCV for video synthesis, set the output frame rate to 30 FPS to ensure smooth video. Select the MP4 video format for output to ensure compatibility and quality. Use the optimized high-resolution images to generate point cloud data, and select the Open3D library for point cloud processing. Use the create_from_rgbd_image() function to combine the high-resolution images with depth information to generate point clouds. Ensure that the density of the point clouds is high enough to capture the details of buildings and vegetation, and set the voxel size of the point clouds to 0.01 meters to improve the fineness of the model. Perform meshing on the point clouds and use the Poisson Surface Reconstruction method to generate a three-dimensional mesh model. Set the reconstruction depth to 9 to ensure the details and smoothness of the mesh. Perform detail optimization on the generated three-dimensional model and apply texture mapping to improve the visual effect of the model and ensure the realism of buildings and vegetation. Export the optimized three-dimensional real scene model to a common 3D file format (such as OBJ or FBX) for use and display on other platforms. Ensure that texture and material information are included during export to maintain the integrity of the model.Load the exported 3D model using 3D visualization software (such as Blender or Sketchfab) for display and interaction, ensuring that the model performs well from different perspectives. Record the visualization effects and generate demonstration videos or screenshots for subsequent presentations or reports.
[0048] In this embodiment, refer to Figure 2 , which is a schematic diagram of the detailed implementation steps of step S1. In this embodiment, the detailed implementation steps of step S1 include:
[0049] Step S11: Obtain the aerial video captured by the drone; perform motion blur convolution suppression processing on the aerial video captured by the drone to construct a motion blur optimized video.
[0050] Step S12: Perform frame-by-frame scene feature point visual recognition on the motion blur optimized video and mark multiple scene feature points in each frame.
[0051] Step S13: Perform dynamic inter-frame correlation matching based on multiple scene feature points in each frame to generate a geometric transformation relationship of the feature points.
[0052] Step S14: Mine the temporal continuity constraints of the geometric transformation relationship of the feature points to generate the temporal continuity constraints between frames.
[0053] Step S15: Perform panoramic geometric stitching on the motion blur optimized video based on the temporal continuity constraints between frames and perform the fusion of the dividing lines at the stitching points to construct a panoramic stitching reconstruction map.
[0054] In this embodiment, a high-performance drone is selected, equipped with a camera with 4K resolution, capable of capturing high-quality images. When selecting the drone, its flight stability and wind resistance are considered to reduce jitter and blur during shooting. The flight mode of the drone is configured as "route flight" to ensure that the drone flies smoothly on the preset route and reduce blur caused by irregular flight. The best shooting time is selected, such as morning or dusk, to obtain good natural lighting conditions and avoid the impact of strong sunlight or shadows on the image quality. The frame rate of the video is set to 30fps, and the resolution is 3840x2160 (4K) to ensure sufficient details and smooth pictures are captured. During recording, ensure that the drone flies at a reasonable speed (such as 2-3 meters per second) to avoid motion blur caused by excessive speed. Use Fourier transform to analyze video frames and identify the motion blur. By converting the image to the frequency domain, observe the loss degree of high-frequency components to judge the direction and intensity of the blur. Set a blur threshold. If the energy of the high-frequency components is lower than the set value (such as 0.1), it is determined that blur exists. Design an inverse motion blur convolution kernel and generate a suitable convolution kernel (such as a 5x5 or 7x7 Box filter) based on the degree of blur. For complex blur situations, use a Gaussian filter and adjust the parameters of the kernel according to the blur direction and degree. Apply the convolution kernel to each frame of the video and use the filter2D function in OpenCV for blur suppression processing. Ensure that the clarity of the image is significantly improved after the convolution operation. Record the image quality metrics (such as PSNR and SSIM) before and after processing to ensure a significant optimization effect. It is expected that the PSNR value increases by at least 10dB and the SSIM value is close to 1 to indicate a significant improvement in image quality. The (Oriented FAST and Rotated BRIEF) algorithm is used for feature point detection. This algorithm is fast and has rotational invariance, suitable for processing dynamic scenes in videos. Configure the parameters of the ORB algorithm, set the number of feature points to 500, and the number of scale pyramid layers to 8 to ensure enough feature points are detected at different scales. At the same time, set Harris corner detection as an auxiliary method to capture key positions. For each frame of the video optimized for motion blur, execute the ORB feature point detection algorithm. In each frame, record the coordinates and descriptors of the detected feature points. Set a minimum brightness threshold to ensure that the detected feature points have sufficient contrast, for example, set it to a 10% brightness change. Draw the detected feature points on each frame of the image using the drawKeypoints function in OpenCV. The feature points are marked as red circles for easy visualization and subsequent analysis. Generate a sequence of marked frames and record the number of marked feature points in each frame to ensure that each frame has at least 10 valid feature points to guarantee the reliability of subsequent matching. When selecting ORB, ensure that the algorithm can effectively process feature points under different lighting and viewing conditions.Select FLANN (Fast Library for Approximate Nearest Neighbors) for feature point matching. FLANN is suitable for handling a large number of feature points and can quickly find the best match. Set the KNN (K-Nearest Neighbors) algorithm with k set to 2 to obtain the two nearest neighbors for each feature point, facilitating subsequent matching accuracy evaluation. Configure the parameters of the FLANN algorithm, set the distance metric (such as Euclidean distance), and set the distance threshold for matching to 30 to ensure that only high-quality matching results are retained. Match the feature points of adjacent frames, compare the feature point descriptors of each frame using the FLANN or KNN algorithm, and record the matching feature point pairs. Set the distance threshold for matching to filter out unmatched pairs that do not meet the conditions, ensuring the accuracy of the matching. Based on the matching feature point pairs, use RANSAC (Random Sample Consensus algorithm) to calculate the geometric transformation relationship (such as homography matrix or fundamental matrix) between each pair of frames. Set the number of iterations to 1000 and the inlier threshold to 3 pixels to improve the robustness of the geometric transformation relationship calculation and ensure the elimination of the interference of outliers on the calculation. Select the Kalman filter as the temporal constraint model. The Kalman filter is suitable for linear systems and can effectively predict the motion state of an object and establish a dynamic model. Through the Kalman filter, the motion trajectory of the feature points can be smoothed to improve the continuity between frames. Set the state transition matrix and observation matrix for the Kalman filter, initialize the state estimate as the initial position of the feature points, and set the covariance matrix to a relatively small value (such as 1.0) to increase the confidence in the initial estimate. According to the geometric transformation relationship between the previous and subsequent frames, use the Kalman filter to calculate the motion trajectory of the feature points and extract the temporal continuity constraint. Record the position changes of each feature point in the time series for subsequent analysis. Iteratively update the temporal data of each feature point to ensure the effectiveness of the continuity constraint. Record the extracted temporal continuity constraints in a data structure, ensuring that each constraint record includes the feature point ID, timestamp, and motion trajectory information. Set a threshold. If the constraint change exceeds the set range (such as 10 pixels), it is marked as abnormal for subsequent processing. Select the multi-view geometry stitching algorithm, which can integrate images from multiple perspectives to generate a panoramic image based on the extracted geometric transformation relationship. Ensure that the temporal continuity constraint of the feature points is considered during the stitching process to enhance the accuracy of the stitching. Set the overlapping area for stitching to 30% to ensure that sufficient feature point matches can be obtained during the stitching process for generating a high-quality panoramic image. Based on the extracted temporal continuity constraints and geometric transformation relationship, perform panoramic geometric stitching on the motion-blurred optimized video. Use OpenCV or a specific stitching library (such as OpenMVG) for processing. During the stitching process, ensure that the overlapping area between each frame and the previous frame can be smoothly transitioned to avoid obvious breaks at the stitching location.Fuse the dividing lines generated during the splicing process using image processing techniques (such as gradient fusion or multiple exposure fusion) to ensure the consistency of illumination and color at the splicing point. Set the size of the fusion area (such as 5 pixels) to ensure a smooth transition at the edge and reduce visible traces at the splicing point.
[0055] In this embodiment, the specific steps of step S11 are as follows:
[0056] Perform histogram adaptive optimization on the UAV aerial video to construct a brightness-optimized aerial video;
[0057] Identify abnormal noise points in the brightness-optimized aerial video and mark the abnormal noise points;
[0058] Filter the abnormal noise points by high-frequency filtering to obtain a noise-filtered video;
[0059] Decompose the noise-filtered video into temporal frames to extract multiple temporal frame images;
[0060] Estimate the optical flow between adjacent frames of multiple temporal frame images to identify the relative jitter data between frames;
[0061] Smooth the inter-frame jitter of multiple temporal frame images according to the relative inter-frame jitter data to construct a smoothed and optimized temporal frame sequence;
[0062] Calculate the optical flow motion of each pixel in the smoothed and optimized temporal frame sequence to extract the pixel motion vectors of each frame image;
[0063] Calculate the motion blur kernel function of the video frame according to the pixel motion vectors of each frame image;
[0064] Perform motion blur convolution suppression processing on the noise-filtered video according to the motion blur kernel function of the video frame to construct a motion blur-optimized video.
[0065] In this embodiment, the CLAHE (Contrast Limited Adaptive Histogram Equalization) method is selected for brightness optimization. Compared with the traditional histogram equalization method, CLAHE can effectively avoid noise amplification caused by over-enhancement. This method enhances local contrast by dividing the image into small regions (called "tiles") and performing equalization processing on each region separately. The image is divided into 8x8 tiles, and the contrast limit is set to 2.0 to ensure that while enhancing the contrast, the influence of noise is suppressed. For each frame of the UAV aerial video, the CLAHE algorithm is applied for histogram equalization. It is implemented using the createCLAHE function in OpenCV. The mean brightness and contrast of the image before and after optimization are recorded to ensure that the brightness and clarity of the processed image are improved, and the mean brightness should be between 100 and 200. Each frame after histogram optimization is recombined into a video file using a video coding format such as H.264 to ensure the quality and compression rationality of the output video. The high-pass filter combined with the threshold method is used to identify abnormal noise points in the image. The high-pass filter can highlight the edges and details in the image, facilitating the identification of noise points. The Laplacian filter is used, and the convolution kernel size is set to 3x3 to enhance the high-frequency information of the image. The Laplacian filter is applied to each frame of the image to obtain the high-frequency image. It is processed using the Laplacian function in OpenCV. A threshold (such as 20) is set to perform binary processing on the high-frequency image to mark the abnormal noise point area. The identified abnormal noise points are marked on the original image using a red box or circle for subsequent processing and analysis. Gaussian filtering is used to smooth the image and reduce the influence of abnormal noise points. The Gaussian filter can effectively remove the high-frequency noise in the image while retaining the edge information. The convolution kernel size of the Gaussian filter is set to 5x5, and the standard deviation is set to 1.0 to ensure a moderate smoothing effect. The Gaussian filter is applied to each frame of the image and implemented using the GaussianBlur function in OpenCV. By processing frame by frame, the influence of abnormal noise points is weakened. Each frame of the filtered image is recombined into a video file to ensure the quality and clarity of the output video. The VideoCapture function in OpenCV is used to extract each frame of the video for subsequent processing. It is set to extract every 1 frame to ensure that sufficient temporal information is captured. The video is traversed using a loop, and each frame is read and saved sequentially. Each frame image can be saved in PNG or JPEG format. The extracted frame images are saved to the specified folder in order with the naming format "frame_001.png" for subsequent processing and identification. The Lucas-Kanade optical flow method is adopted, which can effectively estimate the motion information between adjacent frames and is suitable for processing small-displacement images.Set the window size to 5x5 and the minimum number of feature points to 100 to ensure accurate optical flow estimation. Calculate the optical flow for adjacent frames using the calcOpticalFlowFarneback function in OpenCV. Record the motion vectors of each feature point. Extract the optical flow information of adjacent frames and store the motion vectors of each frame for subsequent analysis and processing. Use the weighted average method to smooth the motion vectors and reduce the motion discontinuity caused by jitter. Set the weight coefficients to 0.7 (current frame) and 0.3 (previous frame) to ensure the smoothing effect on the current frame. Apply the weighted average method to smooth the motion vectors of each frame and eliminate the errors caused by jitter between adjacent frames. Recombine the smoothed frame images to form a new time-series video to ensure the smoothness of the video. Use the Horn-Schunck optical flow method, which is suitable for estimating the motion vectors of each pixel and provides a smooth optical flow field. Set the smoothing parameter to 0.1 to ensure the smoothness and continuity of the optical flow field. Calculate the optical flow for the smoothed and optimized time-series frame sequence, and use the calcOpticalFlowFarneback function in OpenCV to calculate the motion vectors of each pixel. Save the pixel motion vectors of each frame as an array for subsequent analysis and processing. Calculate the motion blur kernel based on the pixel motion vectors, and select the linear blur model, which is suitable for describing the motion direction and speed. Calculate the size and direction of the blur kernel according to the motion vectors, set the kernel size to 5x5, and the blur radius is determined by the magnitude of the motion vectors. Generate the corresponding motion blur kernel function according to the pixel motion vectors of each frame. Use the two-dimensional Gaussian function to describe the shape and intensity of the blur kernel. Save the generated motion blur kernel function for subsequent applications. Use convolution operations to suppress motion blur in the video, combined with the previously generated motion blur kernel function. Ensure that the boundary processing method of the convolution operation is set to "edge padding" to avoid information loss at the boundaries. Apply the motion blur kernel function to each frame of the noise-filtered video for convolution. Implement it using the filter2D function in OpenCV. Recombine each frame of the video after convolution processing into a video file. Ensure that the output video has a significant improvement in image quality and clarity.
[0066] In this embodiment, refer to Figure 3 , which is a schematic diagram of the detailed implementation steps of step S2. In this embodiment, the detailed implementation steps of step S2 include:
[0067] Step S21: Perform per-pixel depth semantic recognition on the panoramic stitching reconstruction map and mark each scene element;
[0068] Step S22: Analyze the temporal displacement changes of each scene element to generate the temporal displacement trajectory of each element;
[0069] Step S23: Classify dynamic and static elements based on the temporal displacement trajectories of each element, so as to obtain dynamic scene elements and static scene elements;
[0070] Step S24: Conduct a full-map spatial distribution analysis on the dynamic scene elements and static scene elements to obtain dynamic and static element spatial distribution data;
[0071] Step S25: Divide the panoramic stitching reconstruction map according to the dynamic and static element spatial distribution data and perform precise boundary segmentation, so as to generate urban building areas and natural environment areas.
[0072] In this embodiment, DeepLabv3+ or Mask R-CNN is selected as the algorithm for deep semantic segmentation. These models perform excellently in pixel-level classification of complex scenes and can handle the recognition of various scene elements. DeepLabv3+ uses dilated convolution to expand the receptive field, which is suitable for processing details in high-resolution images. The model is pre-trained on large datasets (such as Cityscapes or Pascal VOC) to ensure its ability to recognize various elements in urban and natural scenes. The model is fine-tuned and trained according to the dataset of the panoramic stitching reconstruction map, with a learning rate of 0.0001, a batch size of 8, and a training period of 50 epochs. For each frame of the panoramic stitching reconstruction map, the trained model is used for per-pixel deep semantic recognition. Generate the class label of each pixel to identify elements such as buildings, roads, and vegetation in the scene. Set a threshold to ensure that the confidence of each pixel is higher than 0.5 to ensure the accuracy of the marking. Generate a marked image according to the recognition result, using different colors to represent different scene elements. Buildings are represented in red, roads in blue, and vegetation in green. Save the marking information as an image file for subsequent analysis and processing. Extract the position information of each scene element from the marked image obtained in step S21 to generate the initial coordinates and categories (such as buildings, roads, vegetation) of the elements. Set a unique identifier for each element for tracking in subsequent analysis. Integrate the time-series frame data of the panoramic stitching reconstruction map into a list, recording the timestamp of each frame and the corresponding marking information. Ensure that the element information of each frame can correspond to the time series. For each scene element, calculate its displacement change between consecutive frames. Record the displacement information of each element in each frame to generate a time-series displacement trajectory. According to the calculated displacement information, generate a time-series displacement trajectory map of each element. Use different colors or line types to represent different types of scene elements for easy observation of the movement state of the elements. Save the trajectory data as a CSV file for subsequent analysis. Set the classification criteria for dynamic and static elements: if the displacement of an element in the video is greater than 5 pixels, it is determined as a dynamic element; otherwise, it is determined as a static element. Set a time window, such as observing the displacement of an element within 3 frames, to reduce misjudgment. Traverse the time-series displacement trajectory of each scene element and classify it according to the set criteria. If the displacement exceeds the threshold, it is marked as dynamic; otherwise, it is marked as static. Record the classification results, including element ID, category (dynamic or static), and displacement information. Use a two-dimensional histogram or heatmap method to analyze the spatial distribution of dynamic and static elements. Generate a corresponding heatmap according to the coordinates of the elements for intuitive observation of the element distribution. Set the resolution of the heatmap, for example, divide the entire image into a 50x50 grid. Statistically analyze the position information of each dynamic and static scene element and calculate the number of elements in each grid.Generate a heatmap using the Matplotlib library in Python, representing different numbers of elements with different colors to ensure that the heatmap can clearly reflect the distribution of elements. Adopt a region division method based on image segmentation, such as GrabCut or K-means clustering, to facilitate the segmentation of the image into urban building areas and natural environment areas. Set the initial parameters for region division according to the spatial distribution data obtained from the previous analysis. Set the initial masks for the foreground and background in GrabCut and set the number of iterations to 5 to ensure the accuracy of segmentation. Apply the selected region division method to the panoramic stitching reconstruction map to generate mask maps for urban building areas and natural environment areas. Record the number of pixels and the area of each region for subsequent analysis. For the segmented regions, use morphological operations (such as closing) to smooth the boundaries to ensure the accuracy and coherence of the region boundaries. Draw the segmentation boundaries on the segmented image to ensure that each region is clearly visible.
[0073] In this embodiment, step S3 includes the following steps:
[0074] Step S31: Calculate the scene light source intensity based on the UAV aerial video to obtain the scene light source intensity;
[0075] Step S32: Perform environmental light reflection evolution on the urban building area based on the scene light source intensity to generate building environmental light reflection evolution characteristics;
[0076] Step S33: Detect the light shadow projection range of the urban building area and extract the shadow projection range;
[0077] Step S34: Perform dynamic light rendering evolution based on the building environmental light reflection evolution characteristics and the shadow projection range to generate dynamic light rendering evolution characteristics;
[0078] Step S35: Calculate the wall thickness of the urban building area to generate the wall thickness of each building;
[0079] Step S36: Identify the top structure details according to the urban building area and extract the top structure details;
[0080] Step S37: Perform building physical form structure analysis based on the wall thickness of each building and the top structure details to generate the physical form structure data of each building;
[0081] Step S38: Perform 3D building point cloud modeling on the dynamic light rendering evolution characteristics and the physical form structure data of each building to construct a 3D urban building rendering model.
[0082] In this embodiment, in order to calculate the light source intensity of the scene, it is first necessary to perform brightness analysis on the image. By grayscaling the image and applying a pixel-level light intensity detection algorithm (such as a simple image brightness value calculation formula: brightness = 0.299×R + 0.587×G + 0.114×B), the brightness of each frame of the scene is obtained. Based on the brightness values of each frame of the image, image segmentation techniques (such as threshold-based segmentation, K-means clustering, or graph cut algorithms) are applied to extract the light source area from the environment. Through the local contrast and brightness distribution of the image, the identification of the light source area is further optimized, and classification is performed according to the types of light sources in the environment (such as sunlight, reflected light, etc.). Within the detected light source area, a physical lighting model (such as the Phong reflection model or the Lambertian model) is further used to estimate the intensity. Combining the reflection characteristics of the image, the light source intensity of different areas in the scene is obtained. By calculating the direction and intensity of the light source pixel by pixel, it is further mapped in 3D space. According to the light source intensity data obtained in step S31, a physically based light reflection model (such as the Cook-Torrance model) is applied to simulate the light reflection characteristics. Each urban building surface in the scene (such as different materials like glass, metal, concrete, etc.) is modeled through this model, considering the intensity and distribution of the light source reflected by the building surface. According to the scene light source intensity and building materials, the light reflection conditions at different angles and positions are calculated. By calculating the reflection direction and intensity pixel by pixel, the reflection evolution data of each building surface is generated. Specifically, when the light source intensity changes, the reflection intensity, glossiness, etc. of the building surface will change, thereby affecting the visual presentation of the building. Considering the changes of the light source under different times and environmental conditions, a dynamic analysis of the light reflection evolution of the building area is carried out, especially the reflection differences related to day-night changes and weather changes. By taking multiple shots under different lighting conditions and analyzing their changes, the reflection evolution characteristics are further refined. Light reflection model: Cook-Torrance model, the reflection coefficient is set to the typical value of the building surface material (such as the glass reflection coefficient is 0.85, and the concrete is 0.25), Dynamic light source change time periods: Shoot at 8:00, 12:00, and 18:00 every day respectively to simulate the reflection evolution under different lighting conditions, Light intensity range: 500–1500 lux. Use light and geometric information to detect the shadow projection range of the building. First, by calculating the surface normal direction of each building and combining the position and direction of the light source, the projection path of the light is simulated. In this way, the shadow area between the building surface and the ground is obtained. Use a light-based shadow projection algorithm (such as the ray tracing algorithm) to combine the light and the building geometry to simulate the shadow projection effect at different times. Through this algorithm, the shadow range of the building at different times is extracted.Extract the shadow projection range of each building and optimize the shape of the shadow area through image post-processing (such as edge detection and contour extraction) to ensure that the edges of the shadows are smooth and natural. Use the shadow projection map to detect the shadow coverage of each building at a given moment. Assume that the light source height is 50 meters and the light source angle is 30°–90°, and perform shadow detection at different light angles respectively. Combine the building environment light reflection evolution features generated in step S32 and the shadow projection range extracted in S33, and apply a lighting rendering engine (such as Unreal Engine or Unity 3D) for dynamic light rendering. This rendering process will dynamically adjust the appearance and lighting effects of the building based on light source changes, reflection evolution, and shadow projection. Use real-time ray tracing algorithms to perform fine-grained ray simulation on the building. By simulating the reflection and refraction of light on the building surface and in the shadow area, dynamically render the appearance features of the building and generate rendering features related to lighting changes. Analyze the rendering results to optimize the reflection effect of light and shadow details, so that the building can present a natural visual effect under different lighting conditions and enhance the realism of the 3D model. Rendering time: Render 60 frames per second to ensure real-time dynamic effects, and set the number of light tracing paths to 10 to ensure the true presentation of light details. Through image processing algorithms (such as edge detection and depth map generation), extract the external contour and wall structure of the building. Combine the 3D point cloud data of the building and apply deep learning algorithms (such as Mask R-CNN) to identify the edges of the building walls. Estimate the thickness of the walls based on the structural characteristics of the building (such as building drawings and typical wall thickness standards) and the point cloud data. Use a deep learning regression model to predict the thickness of different building materials (such as brick walls and reinforced concrete walls). Verify through actual measurement data and the standard values of wall thickness to ensure the accuracy of the wall thickness estimation results for each building. Wall thickness standard: According to the building type, the wall thickness is estimated to be between 300mm and 700mm, and the point cloud density is 10 points / cm² to ensure accurate extraction of the wall boundary.
[0083] Combined with multi-angle images captured by drones, apply deep learning-based object detection algorithms (such as YOLO, Faster R-CNN) to analyze the building tops, extract the shapes of the roofs, roof facilities (such as skylights, chimneys, etc.), and the texture features of the roofs. Use semantic segmentation algorithms (such as U-Net) to precisely segment the roof areas and extract the specific structural details of each building top, such as facilities like skylights, solar panels, chimneys, etc. Through these detail identifications, improve the accuracy of the roof structure and ensure the realism of subsequent modeling. Apply the extracted top structural details to 3D modeling software to form detailed roof models. Ensure that these details can present high-precision effects from different perspectives. Obtain information on all angles of the roof through multiple shootings (with an interval of 10 degrees each time), and based on the wall thickness and roof structure detail data, construct a 3D physical form model of the building. First, combined with the wall thickness and top structural details, use computer-aided design (CAD) tools to model the building. Through a detailed analysis of the physical form of the building (such as wall strength, roof load-bearing, etc.), perform necessary optimizations on the model to ensure that the physical structure of the building can meet the actual application requirements. Generate physical form data for each building, including features such as the size of the building, wall materials, roof facilities, etc. The physical form data of each building includes parameters such as height, width, depth, wall thickness, roof type, building materials, etc.
[0084] According to the dynamic light rendering features in S34 and the physical form data in S37, use a point cloud generation algorithm (such as the Poisson reconstruction algorithm) to combine the geometric shape and light reflection characteristics of the building to generate a high-precision 3D building point cloud. In 3D modeling software, combine the point cloud data and material textures for rendering processing of the building. Particularly consider the influence of light and shadow to ensure that the appearance of the building is natural and realistic under different lighting conditions. Integrate all the building volume, detail, and lighting data to generate a complete 3D urban building rendering model and perform performance optimization to make it suitable for application scenarios such as virtual reality and augmented reality. Point cloud density: The point cloud density is 100,000 points / m² to ensure the accurate presentation of building details. Rendering accuracy: The texture resolution is 4096x4096 to ensure high-quality rendering effects.
[0085] In this embodiment, step S4 includes the following steps:
[0086] Step S41: Conduct fine-grained identification of vegetation types in the natural environment area to generate multi-region vegetation types;
[0087] Step S42: Estimate the vegetation coverage area of the multi-region vegetation types to obtain the coverage area of each vegetation type;
[0088] Step S43: Perform stereo vision analysis on the natural environment area to generate regional stereo vision information;
[0089] Step S44: Analyze the terrain undulation characteristics based on the regional stereo vision information to obtain the terrain characteristics of the environmental area;
[0090] Step S45: Fit the three-dimensional vegetation distribution to the terrain characteristics of the environmental area according to the coverage area of each vegetation type, and construct a three-dimensional vegetation area model;
[0091] Step S46: Perform three-dimensional panoramic fusion on the three-dimensional urban building rendering model and the three-dimensional vegetation area model to construct a three-dimensional panoramic model.
[0092] In this embodiment, image data of natural environment areas is extracted from UAV aerial video or image data. Since natural environment areas usually include various vegetation types, such as grasslands, shrubs, trees, etc., different types of vegetation have significant differences in terms of color, texture, morphology, etc. Deep learning algorithms (such as deep convolutional neural network CNN, or networks specifically designed for image segmentation such as U-Net) are used to perform fine-grained classification on the vegetation areas. During the training process, different categories of vegetation are used for annotation, such as grasslands, trees, shrubs, etc., and a fine-grained regional division is achieved with the help of a multi-level feature extraction model. This process relies on high-resolution images and sufficient training datasets, and a pre-trained model is used to extract and identify the specific type of each vegetation unit in the image. According to the prediction results of the model, the natural environment area is subdivided into multiple vegetation type areas. Each area corresponds to a specific vegetation type, and the location information and boundaries of each area are generated. Further, geographic information system (GIS) technology is combined to accurately obtain the spatial distribution of the vegetation. Using the spatial distribution data of each vegetation type in the image, based on the pixel-level segmentation results, the area occupied by each vegetation type is calculated. By converting the size of each pixel into the actual geographical area, the coverage area of each vegetation type in the natural environment area is obtained. To ensure high-precision area calculation, GPS information in remote sensing images and the image coordinate system are combined for spatial geometric transformation and geographic coordinate calibration. Especially under different image angles, different times or different geographical environments, image registration techniques (such as the SURF algorithm) are used for image alignment to ensure the consistency of area calculation. According to the image data at multiple times, different weather and lighting conditions, the coverage area of each vegetation type is comprehensively calculated. Through multiple aerial surveys and image acquisitions, the spatio-temporal accuracy of the data is ensured, and calibration is performed in a dynamically changing environment. Using the binocular camera or multi-view image data of the UAV, a depth map of the area is generated through a stereo matching algorithm (such as Semi-Global Matching, SGM). This depth map contains the relative distance between each pixel point and the camera, thereby constructing the stereo vision information of the natural environment area. First, the image quality is improved through image preprocessing techniques (such as light compensation and denoising), then the binocular images are aligned, and different perspectives in the same scene are found through pixel point matching. Using the principle of triangulation, the depth information of each point is calculated, and a stereo image with depth values is generated. The multi-view images are used for depth optimization, and during the reconstruction process, a smoothing algorithm and a depth correction model are combined to remove noise points and mis-matched areas, ensuring that the generated stereo vision information has higher accuracy. Using the depth information obtained through stereo vision analysis, terrain analysis algorithms (such as slope analysis, undulation analysis) are used to perform terrain undulation analysis on the natural environment area. Slope analysis calculates the slope of the ground surface based on the elevation difference of each pixel point, and then determines the undulation of the terrain.Combine the elevation data in the depth map to extract different features of the terrain, such as flat, hill, valley and other areas. Use Geographic Information System (GIS) tools for terrain classification, automatically label the terrain type of each area, and calculate indicators such as the area and slope of each type of terrain. Combine different vegetation types with terrain features for data fusion and modeling to generate multi-dimensional terrain feature data. These data provide a basis for subsequent vegetation distribution modeling and three-dimensional landscape reconstruction. Based on the vegetation type coverage area data and terrain features obtained in the previous steps, use three-dimensional modeling algorithms (such as three-dimensional interpolation method based on particle swarm optimization) to fit each vegetation type on the terrain surface. Through this fitting method, the distribution of different vegetation types is accurately fitted to the terrain surface, taking into account the influence of different terrains, and the three-dimensional spatial distribution of each vegetation type within the region is generated. During the three-dimensional modeling process, it is necessary to fit the three-dimensional spatial distribution of the vegetation according to the characteristics of each vegetation type (such as the height of the trees, the distribution density of the shrubs, etc.). Use the vegetation model to be coupled with the terrain model to ensure that the plant distribution matches the undulating characteristics of the terrain. Combine the existing vegetation data and ground survey data to optimize the three-dimensional vegetation model to ensure the authenticity and accuracy of the vegetation distribution. By comparing different vegetation distribution models, select the optimal fitting scheme. Vegetation distribution density: Set the distribution density according to the vegetation type, the tree distribution density is 1-3 trees per square meter, and the shrub is 5-10 plants per square meter. Use the particle swarm optimization algorithm for distribution fitting, with the error controlled within ±0.2 meters. Integrate the three-dimensional urban building rendering model in step S38 with the three-dimensional vegetation area model in step S45. First, ensure that the spatial coordinates of the two are consistent. Through coordinate transformation and alignment techniques (such as the ICP algorithm), the two models are merged into a unified three-dimensional panoramic model. During the model integration process, use physical rendering techniques such as lighting, shadows, and reflections (such as ray tracing, ambient occlusion) to enhance the interaction effect between the building and the vegetation, making it show more natural light and shadow changes. At the same time, optimize the visual effect of the panoramic model through the interaction between the growth state of the vegetation and the lighting. Finally, output a complete three-dimensional panoramic model, which is applicable to application scenarios such as virtual reality, augmented reality, or digital city. Ensure that the model can be smoothly rendered at different resolutions and guarantee high-quality visual effects. Rendering accuracy: Use high-quality ray tracing rendering, with a resolution of 3840×2160. Fusion algorithm: Use the ICP (Iterative Closest Point) algorithm for model alignment, with the error controlled within 2 cm. Rendering time: The rendering time for each scene is controlled within 30 seconds.
[0093] In this embodiment, the specific steps of step S5 are as follows:
[0094] Step S51: Detect the blurred area of the three-dimensional panoramic model and mark the blurred area of the model;
[0095] Step S52: Perform three-dimensional spatial positioning on the model's blurred area and extract the three-dimensional spatial coordinates of the blurred area;
[0096] Step S53: Based on the three-dimensional spatial coordinates of the blurred area, use a drone to conduct a second low-altitude aerial photography and extract the low-altitude high-definition video of the blurred area.
[0097] In this embodiment, relevant data is extracted from the 3D panoramic model, including information such as geometric shapes, textures, and lighting. Since some areas in the 3D model are blurred due to modeling accuracy, image quality, or perspective issues, we need to detect these areas. To detect the blurred areas, image recognition methods based on deep learning or image quality analysis techniques are adopted. A deep learning model (such as a convolutional neural network CNN) is trained to recognize blurred image areas. In this process, the model analyzes the clarity of each face or texture area in the 3D model. By comparing the clarity of the texture and the complexity of geometric details, the areas with relatively blurred textures and unclear details are marked. Common image quality evaluation metrics such as the Structural Similarity Index (SSIM) or the Peak Signal-to-Noise Ratio (PSNR) are used to quantify the blur degree of the image or model area. The lower the SSIM or the lower the PSNR, the higher the blur degree of the area. The panoramic model is scanned using a blurred area detection algorithm to determine which areas do not meet the clarity standard. For these areas, marking information including the position, area size, and blur degree is added to the 3D model. The SSIM value and the PSNR value are used to evaluate the area clarity. The SSIM threshold is set to 0.7, the PSNR threshold is 25 dB, the model resolution: the resolution of the 3D panoramic model is 4096×4096, and the accuracy is controlled within 0.5 meters. The calculation method: Convolutional Neural Network (CNN) is used for blur detection. The training dataset contains 10,000 model images with different clarity levels. Once the blurred areas are identified through image processing and the deep learning model, the next step is to determine the spatial position of the blurred areas through 3D geometric modeling techniques. This process requires locating the blurred areas in a 3D coordinate system. 3D reconstruction algorithms (such as the least squares method, 3D reconstruction algorithm, or rigid transformation algorithm) are used to extract the spatial coordinates of the detected blurred areas. At this time, the position of the blurred area will be converted from the 2D projection of the model to a 3D spatial coordinate system. Combining the data structure of the 3D model, these blurred areas are located through point cloud data or mesh models. Specifically, the exact 3D coordinates of each marked blurred area are calculated through interpolation techniques. Technologies such as triangular meshes or octrees are used to ensure the accuracy of coordinate extraction. To improve the positioning accuracy, panoramic images are compared with ground actual measurement data to verify the accuracy of the coordinates. If necessary, further geometric correction is performed to reduce errors. After obtaining the 3D spatial coordinates of the blurred areas, the next step is to conduct low-altitude secondary aerial photography using a drone. First, according to the 3D coordinates of the blurred areas, the flight path of the drone is planned. Ensure that when the drone is flying, the camera can accurately cover the blurred areas. This process needs to be planned through flight planning software (such as Mission Planner), and parameters such as flight altitude, shooting angle, and speed are controlled. A high-resolution camera device (such as a 4K high-definition camera or a high-performance thermal imaging camera) is selected to obtain clearer and more detailed video data.When flying, the drone needs to accurately reach and photograph the target area according to the real-time obtained three-dimensional coordinates and GPS information, especially between complex terrains or buildings, to ensure that sufficient clear image data can be obtained in the blurred area. During the aerial photography process, when extracting high-definition video data, high-speed imaging technology is adopted to ensure the smoothness and clarity of the video. Multiple perspectives of shooting with multiple lenses are selected to further improve the video quality. In order to ensure effective blurred area data, each frame in the video needs to be screened and optimized through image processing algorithms to avoid interference from irrelevant areas.
[0098] The collected low-altitude high-definition video needs to be transmitted and stored in the data server in a timely manner. During the storage process, data compression and encryption technologies are applied to ensure the integrity and security of the video data. Then the video data is imported into the subsequent image processing system for analysis and optimization.
[0099] In this embodiment, the specific steps of step S6 are as follows:
[0100] Step S61: Analyze the texture details of the low-altitude high-definition video of the blurred area to generate the texture details of the blurred area;
[0101] Step S62: Identify the geometric shape of the low-altitude high-definition video of the blurred area and extract the geometric shape features of the blurred area;
[0102] Step S63: Optimize the local details of the blurred area of the three-dimensional panoramic model based on the texture details of the blurred area and the geometric shape features of the blurred area, so as to construct a three-dimensional real-scene optimized model.
[0103] In this embodiment, high-definition video data captured by a drone at low altitude is collected. To improve the accuracy of texture detail analysis, necessary video data preprocessing is carried out, including denoising, illumination correction, image enhancement, etc. These operations ensure that the texture details in the video are not interfered by noise and avoid false detection caused by illumination changes. Texture analysis techniques in computer vision (such as Gray-Level Co-Occurrence Matrix (GLCM), Local Binary Pattern (LBP)) or deep learning methods such as Convolutional Neural Network (CNN) are used to extract texture information from video frames. By analyzing local regions of the image, features of the texture (such as contrast, uniformity, roughness, etc.) are extracted, and these features are crucial for judging the detail level of the region. For blurred regions, deep learning algorithms are used to supplement and enhance texture details. An image enhancement model based on Generative Adversarial Network (GAN) is used to generate details close to real textures by learning high-quality texture samples in the training dataset, filling in the missing parts of the original blurred regions. The texture information extracted from each frame of the video image is integrated to generate high-definition texture details of the blurred regions. These texture details will be marked and recorded and prepared for the optimization of subsequent models. Key frames of the blurred regions are extracted from the low-altitude high-definition video and geometric shape analysis is carried out. Through video frame processing, motion blur, noise, and illumination differences in low-altitude shooting are eliminated. To extract the geometric shape features of the blurred regions, traditional computer vision methods such as Hough transform, edge detection (such as Canny operator), and image segmentation techniques are used to identify shapes in the blurred regions. Through edge extraction and contour analysis, basic geometric features of the region, such as lines, curves, angles, and surfaces, are obtained. For the identified geometric shapes, morphological analysis techniques (such as dilation, erosion, opening, and closing operations) are used to further optimize the geometric structure and enhance the geometric features of the region. Convolutional Neural Network (CNN) in deep learning is used for geometric shape classification and recognition. By learning the features of different geometric shapes from the training dataset, geometric details of complex structures are further extracted. Through the above steps, geometric feature information of the blurred regions is extracted, including the basic shape, size, proportion, contour, etc. of the object. And these information are further transformed into three-dimensional space coordinates and prepared for the optimization of the subsequent three-dimensional model. In the three-dimensional panoramic model, the spatial coordinates and geometric shape features of the blurred regions are marked. First, the texture features obtained through texture detail analysis (step S61) are combined with the geometric structure information obtained through geometric shape recognition (step S62) to identify the local regions that need to be optimized. For the blurred regions, through texture mapping technology, the optimized texture details are extracted from the low-altitude high-definition video and mapped to the blurred regions of the three-dimensional panoramic model. High-resolution texture images are used to replace the original blurred textures to ensure that the texture details of the optimized regions are more abundant and clear. For the geometric shape features, through geometric reconstruction techniques (such as point cloud-based reconstruction methods or surface fitting techniques), the geometric shapes of the blurred regions are refined to make them more conform to the structure of the real scene.Combining the texture and geometric information of the fuzzy region, use 3D model fusion techniques (such as B-spline interpolation or the finite element method) to perform smooth transitions on the local area, ensuring that the optimized details are naturally integrated into the existing 3D panoramic model without abrupt changes. For dynamic features such as lighting, shadows, and reflections, use ray tracing technology to optimize the details of local lighting and reflections. After the local detail optimization is completed, perform a consistency check on the entire 3D model to ensure that the optimized texture and geometric features are consistent with the style of the original panoramic model. Finally, generate the final 3D real-scene optimized model for application scenarios such as virtual reality, augmented reality, and digital cities.
[0104] In this embodiment, a drone-based real-scene 3D automatic modeling system is provided for performing the drone-based real-scene 3D automatic modeling method described above, including:
[0105] A geometric stitching module for obtaining drone aerial videos; performing motion blur convolution suppression processing and panoramic geometric stitching on the drone aerial videos to construct a panoramic stitching reconstruction diagram;
[0106] A depth semantic recognition module for performing per-pixel depth semantic recognition on the panoramic stitching reconstruction diagram to generate urban building areas and natural environment areas;
[0107] A light rendering evolution module for performing dynamic light rendering evolution and building physical form structure analysis on urban building areas to construct a 3D urban building rendering model;
[0108] A 3D panoramic fusion module for performing fine-grained identification of vegetation types in natural environment areas and performing 3D panoramic fusion based on the 3D urban building rendering model to construct a 3D panoramic model;
[0109] A fuzzy region detection module for detecting fuzzy regions in the 3D panoramic model and performing low-altitude secondary aerial photography to extract low-altitude high-definition videos of fuzzy regions;
[0110] A local detail optimization module for performing local detail optimization of fuzzy regions based on the low-altitude high-definition videos of fuzzy regions to construct a 3D real-scene optimized model.
[0111] The present invention can effectively remove the blurring effect and restore image details by using motion blur convolution suppression algorithms (such as blind deblurring algorithms), ensuring that each frame of the image has a high-quality visual effect. This is particularly important for subsequent 3D modeling, as any degradation in image quality will affect feature extraction and model accuracy. Through panoramic geometric stitching technology, aerial images from multiple perspectives can be accurately stitched into a complete panoramic image. Geometric stitching minimizes the perspective distortion between different images and effectively synthesizes a seamless large-scale view, providing a high-precision image basis for subsequent depth semantic recognition and 3D reconstruction. Through deep learning networks (such as convolutional neural networks, Transformers, etc.), pixel-by-pixel depth semantic recognition can be performed to classify and label each element in the image at the pixel level. This not only helps to accurately identify elements such as buildings, roads, and vegetation, but also can distinguish the boundaries between natural environment areas and urban building areas based on environmental complexity, thereby improving the subsequent modeling accuracy. Physical lighting and dynamic rendering algorithms (such as ray tracing or physically based rendering PBR technology) are used to simulate the light changes at different times and angles. The impact of light on buildings, such as light reflection and shadow changes, is accurately reproduced, providing a sense of realism for the visual presentation of buildings under different lighting conditions. By simulating the evolution of light, the final building model appears more vivid and natural in the real environment. Analyzing the physical structure of the building (such as wall thickness, window frames, roof structure, etc.) can construct a more accurate 3D building model, avoiding the limitations of models based solely on flat drawings. This not only improves the modeling accuracy but also enhances the realism of the building, providing high-quality support for urban planning and virtual reality applications. Through deep learning algorithms (such as FCN, Mask R-CNN, etc.), fine-grained classification and recognition of vegetation types in the natural environment can be carried out. This process can be accurate to different types of plants, trees, etc., and can even identify the growth status of plants (such as healthy, withered, etc.), providing strong data support for the accurate modeling of the natural environment. Combining the modeling results of urban building areas and natural environment areas, the two are integrated into a whole through 3D panoramic fusion technology (such as seamless fusion of lighting, materials, and geometric models). In this way, the final constructed 3D panoramic model not only shows the architectural appearance of the city but also accurately reproduces the details of the surrounding natural environment, enhancing the realism and immersion of the entire scene. Through image processing technology, it is possible to identify the blurred areas caused by distance, movement, etc. during the aerial photography process. This detection step is very important because the blurred areas will affect the accuracy and integrity of the modeling. Timely discovery and marking of these areas provide guidance for subsequent optimization. Once a blurred area is detected, low-altitude secondary aerial photography can be carried out for high-definition acquisition at a closer distance and a more stable flight state.In this way, clearer and higher-resolution video data can be obtained, compensating for the loss of details in the first aerial survey and providing a high-quality data source for the subsequent detail repair of 3D modeling. Through the deep learning-based detail recognition of low-altitude videos, important information such as texture details and structural features in the blurred areas are extracted and applied to the optimization of 3D models. This not only improves the accuracy of the blurred areas but also increases the fineness and realism of the entire 3D model. Based on the optimized detail information, a 3D real-scene optimization model with high precision and quality is constructed. The 3D model after detail repair has higher visual realism and spatial restoration degree and can be applied to multiple fields such as urban planning, virtual reality, and architectural design.
[0112] Therefore, in any regard, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the application documents are intended to be encompassed within the present invention.
[0113] As described above, these are only the specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein but will conform to the widest scope consistent with the principles and novel features invented herein.
Claims
1. A method for real - scene three - dimensional automated modeling based on drones, characterized in that, It includes the following steps: Step S1: Obtain the aerial video captured by the drone; perform motion blur convolution suppression processing and panoramic geometric stitching on the aerial video captured by the drone, so as to construct a panoramic stitching reconstruction map; Step S2: Perform per-pixel depth semantic recognition on the panoramic stitching reconstruction map, so as to generate urban building areas and natural environment areas; Step S3: Perform dynamic light rendering evolution and building physical form structure analysis on the urban building areas, and construct a three-dimensional urban building rendering model; Step S4: Perform fine-grained identification of vegetation types on the natural environment areas, and perform three-dimensional panoramic fusion according to the three-dimensional urban building rendering model, so as to construct a three-dimensional panoramic model; Step S5: Detect the blurred areas in the three-dimensional panoramic model, and perform secondary low-altitude aerial photography to extract the low-altitude high-definition video of the blurred areas; Step S6: Perform local detail optimization of the blurred areas according to the low-altitude high-definition video of the blurred areas, so as to construct a three-dimensional real-scene optimization model; Among them, the specific steps of Step S1 are: Step S11: Obtain the aerial video captured by the drone; perform motion blur convolution suppression processing on the aerial video captured by the drone, so as to construct a motion blur optimized video; Step S12: Perform per-frame scene feature point visual recognition on the motion blur optimized video, and mark multiple scene feature points in each frame; Step S13: Perform dynamic inter-frame correlation matching according to multiple scene feature points in each frame, and generate a geometric transformation relationship of the feature points; Step S14: Mine the temporal continuity constraints of the geometric transformation relationship of the feature points, and generate the temporal continuity constraints between frames; Step S15: Perform panoramic geometric stitching on the motion blur optimized video based on the temporal continuity constraints between frames, and perform segmentation line fusion at the stitching positions, so as to construct a panoramic stitching reconstruction map; Among them, the specific steps of Step S11 are: Perform histogram adaptive optimization on the aerial video captured by the drone to construct a brightness optimized aerial video; Identify abnormal noise points in the brightness optimized aerial video and mark the abnormal noise points; Perform high-frequency filtering on the abnormal noise points to obtain a noise filtered video; Perform temporal frame decomposition on the noise filtered video to extract multiple temporal frame images; Estimate the optical flow between adjacent frames of the multiple temporal frame images to identify the relative jitter data between frames; Smooth the inter-frame jitter of the multiple temporal frame images according to the relative jitter data between frames, so as to construct a smoothed optimized temporal frame sequence; Calculate the pixel motion vector of each frame image by performing per-pixel optical flow motion calculation on the smoothed optimized temporal frame sequence; Calculate the motion blur kernel function of the video image according to the pixel motion vector of each frame image; Perform motion blur convolution suppression processing on the noise filtered video according to the motion blur kernel function of the video image, so as to construct a motion blur optimized video.
2. The method for drone-based real-scene three-dimensional automatic modeling according to claim 1, wherein The specific steps of Step S2 are: Step S21: Perform per-pixel depth semantic recognition on the panoramic stitching reconstruction map and mark each scene element; Step S22: Perform temporal displacement change analysis on each scene element to generate the temporal displacement trajectory of each element; Step S23: Perform dynamic and static element classification according to the temporal displacement trajectory of each element, so as to obtain dynamic scene elements and static scene elements; Step S24: Perform a full-map spatial distribution analysis on dynamic and static scene elements to obtain dynamic and static element spatial distribution data; Step S25: Divide the panoramic stitching reconstruction map according to the dynamic and static element spatial distribution data and perform precise boundary segmentation to generate urban building areas and natural environment areas.
3. The method according to claim 1, wherein The specific steps of Step S3 are as follows: Step S31: Calculate the scene light source intensity based on the UAV aerial video to obtain the scene light source intensity; Step S32: Perform environmental light reflection evolution on the urban building area based on the scene light source intensity to generate building environmental light reflection evolution characteristics; Step S33: Detect the light shadow projection range of the urban building area and extract the shadow projection range; Step S34: Perform dynamic light rendering evolution based on the building environmental light reflection evolution characteristics and the shadow projection range to generate dynamic light rendering evolution characteristics; Step S35: Calculate the wall thickness of the urban building area to generate the wall thickness of each building; Step S36: Identify the top structure details according to the urban building area and extract the top structure details; Step S37: Perform building physical form structure analysis based on the wall thickness of each building and the top structure details to generate the physical form structure data of each building; Step S38: Perform 3D building point cloud modeling on the dynamic light rendering evolution characteristics and the physical form structure data of each building to construct a 3D urban building rendering model.
4. The method for drone-based real-scene three-dimensional automatic modeling according to claim 1, wherein The specific steps of Step S4 are as follows: Step S41: Perform fine-grained identification of vegetation types in the natural environment area to generate multi-region vegetation types; Step S42: Calculate the vegetation coverage area of the multi-region vegetation types to obtain the coverage area of each vegetation type; Step S43: Perform stereo vision analysis on the natural environment area to generate regional stereo vision information; Step S44: Analyze the terrain undulation characteristics according to the regional stereo vision information to obtain the terrain characteristics of the environmental area; Step S45: Perform 3D vegetation distribution fitting on the terrain characteristics of the environmental area according to the coverage area of each vegetation type to construct a 3D vegetation area model; Step S46: Perform 3D panoramic fusion on the 3D urban building rendering model and the 3D vegetation area model to construct a 3D panoramic model.
5. The method for drone-based real-scene three-dimensional automatic modeling according to claim 1, wherein, The specific steps of Step S5 are as follows: Step S51: Detect the blurred areas of the 3D panoramic model and mark the blurred areas of the model; Step S52: Perform 3D spatial positioning on the blurred areas of the model and extract the 3D spatial coordinates of the blurred areas; Step S53: Use the UAV to perform low-altitude secondary aerial photography based on the 3D spatial coordinates of the blurred areas and extract the low-altitude high-definition video of the blurred areas.
6. The method for drone-based real-scene three-dimensional automated modeling according to claim 1, characterized in that The specific steps of Step S6 are as follows: Step S61: Perform texture detail analysis on the low-altitude high-definition video of the blurred areas to generate texture details of the blurred areas; Step S62: Identify the geometric shapes of the low-altitude high-definition video of the blurred areas and extract the geometric shape characteristics of the blurred areas; Step S63: Perform local detail optimization of the blurred areas on the 3D panoramic model based on the texture details of the blurred areas and the geometric shape characteristics of the blurred areas to construct a 3D real-scene optimized model.
7. An unmanned aerial vehicle-based real-scene three-dimensional automatic modeling system, characterized in that, For implementing the drone-based real-scene three-dimensional automated modeling method as described in claim 1, including: A geometric stitching module, configured to obtain aerial videos captured by the drone; perform motion blur convolution suppression processing and panoramic geometric stitching on the aerial videos captured by the drone, so as to construct a panoramic stitching reconstruction map; A depth semantic recognition module, configured to perform pixel-by-pixel depth semantic recognition on the panoramic stitching reconstruction map, so as to generate an urban building area and a natural environment area; A light rendering evolution module, configured to perform dynamic light rendering evolution and building physical form structure analysis on the urban building area, and construct a three-dimensional urban building rendering model; A three-dimensional panoramic fusion module, configured to perform fine-grained recognition of vegetation types in the natural environment area, and perform three-dimensional panoramic fusion according to the three-dimensional urban building rendering model, so as to construct a three-dimensional panoramic model; A fuzzy area detection module, configured to detect fuzzy areas in the three-dimensional panoramic model, and perform secondary low-altitude aerial photography to extract low-altitude high-definition videos of the fuzzy areas; A local detail optimization module, configured to perform local detail optimization of the fuzzy areas according to the low-altitude high-definition videos of the fuzzy areas, so as to construct a three-dimensional real-scene optimization model.
Citation Information
Patent Citations
Space observation and reconstruction method and system based on unmanned aerial vehicle
CN116957360A
Data acquisition and modeling method and system suitable for wide-area surface mine
CN117723029A
Three-dimensional geographic information model rendering method and system
CN118071953A