Method and system for evaluating the precision of real estate 3D modeling based on unmanned aerial vehicle oblique photography
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-28
- Publication Date
- 2026-08-11
AI Technical Summary
这些问题导致其难以满足不动产三维建模对客观性、高精度和可靠性的实际需求
[0012](1)本发明通过将无人机倾斜摄影获取的不动产二维图像与元数据文件中的深度尺度因子、不动产实测尺寸、产权边界坐标以及相机内参和POS地理数据直接输入端到端深度学习流程,先经倾斜摄影校正增强处理和航带偏差修正得到适配模型的第一处理图像与第二处理图像,再在深度估计阶段引入航高一致性约束和建筑结构约束生成物理深度图与可视化深度图,最后将这些结果输入3D视频生成模型并基于物理深度图与实测数据进行空间尺度信息和边界位置的自动比对,从而实现了完全客观、不依赖专家主观判断的建模精度评估,从根本上解决了现有模糊综合评估方法指标权重依赖人工经验、隶属度函数无法自适应不同不动产建筑结构的问题,显著提升了评估结果在住宅、商用等各类不动产场景下的客观性和准确性。
Smart Images

Figure CN122312925B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of real estate modeling and analysis technology, and more specifically, to a method and system for evaluating the accuracy of 3D real estate modeling based on UAV oblique photography. Background Technology
[0002] UAV oblique photogrammetry, with its rapid data acquisition, multi-view coverage, and cost-effectiveness, has become one of the core methods for 3D modeling of real estate. This technology uses UAVs equipped with cameras to acquire oblique image data and integrates it with geographic location information, enabling the efficient construction of real-world 3D models of real estate. It plays a crucial role in areas such as real estate registration and ownership confirmation, digital twin city construction, virtual reality simulation, and geographic information system integration.
[0003] The accuracy of 3D modeling of real estate directly affects the accuracy of property boundary identification, spatial scale measurement, and the reliability of subsequent information management. Therefore, developing scientific and effective accuracy assessment methods for real estate scenarios is of great practical significance.
[0004] In existing technologies, the fuzzy comprehensive evaluation method is mainly used for the accuracy assessment of 3D real estate modeling based on UAV oblique photogrammetry. This method generally first constructs a two-level indicator system including data acquisition and data processing stages, then uses the analytic hierarchy process (AHP) to determine the weight coefficients of each indicator, and finally constructs a fuzzy relation matrix and applies the maximum membership principle to comprehensively calculate the final evaluation result.
[0005] However, existing fuzzy comprehensive evaluation methods have several limitations in real estate application scenarios. On the one hand, the determination of indicator weights relies heavily on the subjective experience of experts, making it difficult to objectively incorporate key elements unique to real estate, such as property boundary coordinates and measured dimensions; furthermore, membership functions require manual design and cannot automatically adapt to structural differences in different types of real estate buildings or actual changes in complex terrain. On the other hand, existing methods mainly handle linear relationships between indicators, with limited ability to capture nonlinear interactions between image data and metadata, as well as building structural features. In addition, the entire evaluation process is time-consuming and has poor applicability in new real estate scenarios. These problems make it difficult to meet the actual requirements of objectivity, high accuracy, and reliability for real estate 3D modeling.
[0006] No effective solutions have yet been proposed to address the problems in the relevant technologies. Summary of the Invention
[0007] To address the problems in related technologies, this invention proposes a method and system for evaluating the accuracy of 3D modeling of real estate based on UAV oblique photography, in order to overcome the aforementioned technical problems existing in the existing related technologies.
[0008] Therefore, the specific technical solution adopted by the present invention is as follows:
[0009] According to one aspect of the present invention, a method for evaluating the accuracy of 3D real estate modeling based on UAV oblique photogrammetry is provided. This method includes: extracting the image resolution and number of channels from a two-dimensional image and metadata file of the real estate; extracting a depth scale factor, measured dimensions of the real estate, property boundary coordinates, and camera intrinsic parameters from the metadata file; spatially mapping the property boundary coordinates to the current image resolution; performing oblique photogrammetry correction and enhancement processing on the two-dimensional image based on the two-dimensional image and the corresponding camera intrinsic parameters, and performing flight strip deviation correction using UAV POS geographic data to obtain a first processed image adapted to a depth estimation model and a second processed image adapted to a 3D video generation model; performing depth estimation with geospatial and structural constraints based on the first processed image, combined with the depth scale factor, UAV POS geographic data, and camera intrinsic parameters to obtain a visual depth map and a physical depth map; inputting the second processed image and the visual depth map into a 3D video generation model to generate a 3D real estate video; and evaluating the modeling accuracy of the generated 3D real estate video based on the physical depth map, measured dimensions of the real estate, and property boundary coordinates, and outputting the evaluation result.
[0010] According to another aspect of the present invention, a real estate 3D modeling accuracy evaluation system based on UAV oblique photogrammetry is also provided. This system includes: an input module for extracting the image resolution and number of channels of the 2D image based on a 2D image and metadata file of the real estate, and extracting the depth scale factor, measured dimensions of the real estate, property boundary coordinates, and camera intrinsic parameters from the metadata file; spatially mapping the property boundary coordinates to the current image resolution; and an image preprocessing module for performing oblique photogrammetry correction and enhancement processing on the 2D image based on the 2D image and the corresponding camera intrinsic parameters, and combining this with UAV POS geographic data. The system performs flight strip deviation correction to obtain a first processed image adapted to the depth estimation model and a second processed image adapted to the 3D video generation model. The depth estimation module is used to perform depth estimation with geospatial and building structure constraints based on the first processed image, combined with depth scale factors, UAV POS geographic data, and camera intrinsic parameters, to obtain a visual depth map and a physical depth map. The 3D video evaluation module is used to input the second processed image and the visual depth map into the 3D video generation model to generate a 3D video of the real estate, and evaluate the modeling accuracy of the generated 3D video of the real estate based on the physical depth map, the measured size of the real estate, and the coordinates of the property boundary, and output the evaluation results.
[0011] The beneficial effects of this invention are as follows:
[0012] (1) This invention directly inputs the two-dimensional images of real estate obtained by UAV oblique photography and the depth scale factor, measured size of real estate, property boundary coordinates, camera intrinsic parameters and POS geographic data in the metadata file into the end-to-end deep learning process. First, the oblique photography correction and enhancement processing and flight strip deviation correction are used to obtain the first processed image and the second processed image of the model. Then, in the depth estimation stage, flight altitude consistency constraints and building structure constraints are introduced to generate physical depth maps and visual depth maps. Finally, these results are input into the 3D video generation model and the spatial scale information and boundary position are automatically compared based on the physical depth map and the measured data. This achieves a completely objective modeling accuracy assessment that does not rely on expert subjective judgment. It fundamentally solves the problems of existing fuzzy comprehensive evaluation methods where the index weights rely on human experience and the membership function cannot adapt to different real estate building structures. It significantly improves the objectivity and accuracy of the evaluation results in various real estate scenarios such as residential and commercial properties.
[0013] (2) While generating 3D videos of real estate, this invention completes the comprehensive judgment of scale assessment accuracy and boundary assessment accuracy based on physical depth maps. The entire process, from image preprocessing to depth estimation after constraint optimization, to diffusion sampling with scale constraints, structural constraints and spatial range constraints, and finally weighted average comprehensive judgment, achieves automated closed-loop processing. Thus, in real estate 3D modeling, it simultaneously ensures the temporal continuity of video output and the quantifiable verification of modeling accuracy. It completely overcomes the limitations of existing technologies, such as weak feature capture capability, low assessment efficiency and difficulty in integrating real estate-specific elements such as property boundary coordinates and measured dimensions. It provides efficient, reliable and directly implementable accuracy guarantee for applications such as real estate registration, digital twin city construction and virtual reality simulation. Attached Figure Description
[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 This is a flowchart illustrating the accuracy assessment method for 3D real estate modeling based on UAV oblique photography according to an embodiment of the present invention.
[0016] Figure 2 This is a schematic diagram of the architecture of the depth estimation model in the real estate 3D modeling accuracy evaluation method based on UAV oblique photography according to an embodiment of the present invention;
[0017] Figure 3This is a schematic diagram of the architecture of the 3D video generation model in the real estate 3D modeling accuracy evaluation method based on UAV oblique photography according to an embodiment of the present invention.
[0018] Figure 4 This is a specific implementation diagram of evaluating the modeling accuracy of generated 3D real estate video in the real estate 3D modeling accuracy evaluation method based on UAV oblique photography according to an embodiment of the present invention;
[0019] Figure 5 This is a principle block diagram of a real estate 3D modeling accuracy assessment system based on UAV oblique photography according to an embodiment of the present invention. Detailed Implementation
[0020] To further illustrate the various embodiments, the present invention provides accompanying drawings, which are part of the disclosure of the present invention. These drawings are mainly used to illustrate the embodiments and can be used in conjunction with the relevant descriptions in the specification to explain the operating principles of the embodiments. With reference to these drawings, those skilled in the art should be able to understand other possible implementation methods and the advantages of the present invention. The components in the drawings are not drawn to scale, and similar component symbols are generally used to represent similar components.
[0021] According to embodiments of the present invention, a method and system for evaluating the accuracy of 3D modeling of real estate based on UAV oblique photography are provided.
[0022] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments, such as... Figure 1 As shown, according to an embodiment of the present invention, a method for evaluating the accuracy of 3D real estate modeling based on UAV oblique photogrammetry is provided. This method includes:
[0023] S1. Based on the two-dimensional image and metadata file of the real estate, extract the image resolution and number of channels of the two-dimensional image, and extract the depth scale factor, actual size of the real estate, property boundary coordinates and camera intrinsic parameters from the metadata file; spatially map the property boundary coordinates to the current image resolution;
[0024] S2. Based on the two-dimensional image and the corresponding camera intrinsic parameters, the two-dimensional image is subjected to oblique photogrammetry correction and enhancement processing, and the flight path deviation is corrected by combining the UAV POS geographic data to obtain the first processed image adapted to the depth estimation model and the second processed image adapted to the 3D video generation model.
[0025] S3. Based on the first processed image, combined with the depth scale factor, UAV POS geographic data and camera intrinsic parameters, perform depth estimation with geospatial constraints and building structure constraints to obtain a visual depth map and a physical depth map.
[0026] S4. Input the second processed image and the visualized depth map into the 3D video generation model to generate a 3D video of the real estate. Based on the physical depth map, the actual measured size of the real estate and the coordinates of the property boundary, evaluate the modeling accuracy of the generated 3D video of the real estate and output the evaluation results.
[0027] In one embodiment, performing oblique photogrammetry correction and enhancement processing on a 2D image, and combining it with UAV POS geographic data for flight strip deviation correction, includes: extracting the image resolution and number of channels based on the 2D image; parsing the camera intrinsic parameters and UAV POS geographic data corresponding to the 2D image; wherein the camera intrinsic parameters include at least focal length, principal point coordinates, and distortion coefficients; the UAV POS geographic data includes at least latitude, longitude, elevation, and flight altitude; converting the 2D image into an RGB three-channel image, and spatially mapping the property boundary coordinates to the current image resolution; performing distortion correction on the converted RGB three-channel image using the camera intrinsic parameters and distortion coefficients; calculating the geographic coordinate offset of the current image relative to the reference image, using the UAV POS geographic data and a reference image within the flight strip; performing flight strip deviation correction on the distortion-corrected image through affine translation transformation based on the geographic coordinate offset to obtain the processed original reference image; and performing size adjustment, normalization, and format conversion on the processed original reference image to obtain a first processed image adapted to the depth estimation model and a second processed image adapted to the 3D video generation model.
[0028] Specifically, building texture enhancement and noise removal are performed on the processed original reference image to obtain a building detail enhanced image; according to the input requirements of the depth estimation model, the building detail enhanced image is resized, normalized, and format converted to obtain a first processed image; according to the input requirements of the 3D video generation model, the processed original reference image is resized, normalized, and format converted to obtain a second processed image.
[0029] In one embodiment, the expression for distortion correction of the converted RGB three-channel image is:
[0030] ;
[0031] In the formula, x d and y d The coordinates of the normalized image after correction; x u and y u represents the normalized image coordinates after distortion; k1, k2, and k3 are the radial distortion coefficients; p1 and p2 are the tangential distortion coefficients;
[0032] The expression for correcting flight strip deviation is:
[0033] ;
[0034] In the formula, u and v are the pixel coordinates of the image after distortion correction; and To complete the image pixel coordinates after flight strip deviation correction; t u and t v These represent the translation amounts of the current image relative to the reference image in the horizontal and vertical directions, respectively.
[0035] In one embodiment, depth estimation with geospatial and building structure constraints, combining depth scale factor, UAV POS geographic data, and camera intrinsic parameters, includes: reading UAV POS geographic data and camera intrinsic parameters corresponding to a first processed image; calculating a depth scale factor based on flight altitude information in the UAV POS geographic data, focal length information in the camera intrinsic parameters, and physical dimensions of image pixels; performing tensor normalization on the first processed image; dividing the first processed image into multiple image patches when the resolution exceeds a preset threshold, and using the image patches as input units of the depth estimation model; using the entire first processed image as the input unit of the depth estimation model when the resolution does not exceed a preset threshold; adding flight strip identification features to input units belonging to the same flight strip, and inputting the input units into the depth estimation model; introducing flight altitude consistency constraints and building structure constraints during the inference process of the depth estimation model to obtain a constraint-optimized relative depth result.
[0036] Specifically, the relative depth results are upsampled and restored, and scaled by combining a depth scale factor to generate a physical depth map; the upsampled and restored relative depth results are normalized and data type converted to generate a visual depth map, so as to obtain depth estimation results for 3D video generation and modeling accuracy evaluation.
[0037] In one embodiment, introducing flight altitude consistency constraints and building structure constraints during the inference process of the depth estimation model to obtain the constraint-optimized relative depth result includes: determining the theoretical reference depth of the ground area based on the flight altitude information in the UAV POS geographic data and the focal length information in the camera intrinsic parameters; constructing flight altitude consistency constraints based on the deviation between the theoretical reference depth and the predicted depth of the ground area output by the depth estimation model to ensure that the predicted depth of the ground area is consistent with the actual flight altitude; applying building structure constraints to the relative depth result output by the depth estimation model based on the edge position, continuous wall area, and corner area of the real estate buildings in the image to suppress depth noise in non-building boundary areas and maintain depth changes at the building outline position; jointly optimizing the loss function after introducing flight altitude consistency constraints and building structure constraints to output the constraint-optimized relative depth result; wherein, the building structure constraints include edge smoothing constraints and local second-order structure constraints.
[0038] In one embodiment, the edge-preserving smoothness constraint is constructed based on edge-aware weights determined by the luminance channel gradient of the original RGB image, combined with the first-order gradients of the depth map in the horizontal and vertical directions, to reduce depth fluctuations within the same building plan area; the local second-order structure constraint is constructed based on the local second-order structure response of the depth map, combined with the edge-aware weights, to improve the structural consistency of building walls, balconies, and corner areas; the expression for calculating the depth scale factor is:
[0039] ;
[0040] In the formula, k is the depth scale factor; H is the relative flight altitude in the UAV POS geographic data; d p f is the physical size of the image pixels; x The horizontal focal length is a camera intrinsic parameter.
[0041] The expression for the altitude consistency constraint is:
[0042] ;
[0043] In the formula, Loss due to flight altitude consistency; This represents the total number of ground pixels. For the ground mask, when the pixel If it belongs to the ground area, use 1; otherwise, use 0. For depth estimation models at pixel points The predicted relative depth at the output; For pixels The theoretical reference depth at that point.
[0044] In one embodiment, generating a 3D video of real estate includes: reading a second processed image and a visualized depth map; acquiring the depth scale factor, measured dimensions of the real estate, UAV POS geographic data, and property rights feature parameters corresponding to the second processed image, and constructing spatial constraint information based on the UAV POS geographic data and the measured dimensions of the real estate; setting camera motion parameters according to a preset real estate camera motion mode, and inputting the second processed image, visualized depth map, spatial constraint information, and camera motion parameters into a 3D video generation model for constrained model inference to obtain a video frame sequence; when the resolution of the second processed image exceeds a preset threshold, performing block generation and fusion processing on the video frame sequence; performing video synthesis and optimization processing on the generated video frame sequence, and outputting a 3D video of real estate.
[0045] In one embodiment, constrained model inference to obtain a video frame sequence includes: establishing a scale correspondence between the generated frame depth map and the actual physical dimensions of the real estate based on the depth scale factor and the measured dimensions of the real estate; introducing the scale correspondence as a scale constraint into the diffusion sampling process of the 3D video generation model to limit the scale variation range of the building body during the generation process; determining the building wall area and ground area based on the visualized depth map, and constructing structural constraints based on the normal relationship between the building wall area and ground area; introducing the structural constraints into the diffusion sampling process to maintain the vertical and horizontal structural relationship between the building wall and the ground; constraining the spatial range of the camera pose corresponding to each frame during the diffusion sampling process based on the measured dimensions of the real estate and the scene center point; and outputting a video frame sequence based on the diffusion sampling results after introducing scale constraints, structural constraints, and spatial range constraints.
[0046] Specifically, the block generation and fusion processing of the video frame sequence includes: when the resolution of the second processed image exceeds a preset threshold, dividing the second processed image into several image sub-blocks as input units for the 3D video generation model; adding spatial location identifiers to each image sub-block, and generating corresponding video sub-frame sequences based on the spatial location identifiers to maintain the positional correspondence of each image sub-block in the overall scene; performing brightness correction and color correction on each video sub-frame sequence to obtain corresponding brightness correction results and color correction results; performing weighted fusion on adjacent video sub-frames after brightness and color correction in the overlapping area to eliminate sub-block splicing boundaries; and obtaining the fused video frame sequence.
[0047] In one embodiment, such as Figure 4 As shown, the modeling accuracy evaluation of the generated 3D real estate video includes: based on the physical depth map, combined with camera intrinsics and UAV POS geographic data, key feature points of the real estate are selected, and the coordinates of each key feature point in three-dimensional space are calculated using back projection of the physical depth values. The spatial scale information of the real estate is determined by calculating the Euclidean distance between the coordinates of each key feature point in three-dimensional space. The spatial scale information includes the length, width, and height of the real estate. The spatial scale information is compared with the corresponding measured length, width, and height in the actual measured dimensions of the real estate, and the relative error percentage of each dimension is calculated. The average of the three is taken as the scale evaluation accuracy. Based on the property boundary coordinates, combined with UAV POS geographic data and camera intrinsics, the property boundary coordinates are projected onto the image coordinate system of the corresponding frame of the real estate 3D video to obtain the reference boundary position. The actual boundary position of the real estate is extracted from the real estate 3D video, and the pixel coordinates of the actual boundary position are compared with the reference boundary position to calculate the average offset distance as the boundary evaluation accuracy. Combining the scale evaluation accuracy and the boundary evaluation accuracy, the modeling accuracy of the real estate 3D video is comprehensively judged, and the corresponding evaluation result is output.
[0048] Specifically, key feature points include building corner points, edge points, and top points. The extraction process involves calculating the depth gradient of the physical depth map to highlight the structural boundaries and corner positions of the main property. Then, the location information of the building corner points, edge points, and top points is automatically identified through threshold segmentation. This extracted location information is then directly mapped back to the corresponding image coordinate system. This allows for subsequent back-projection calculation of the coordinates of each key feature point in three-dimensional space using the physical depth values. By calculating the Euclidean distance between these coordinates, the spatial scale information of the main property, such as its length, width, and height, can be obtained.
[0049] Specifically, the actual boundary position is determined by performing connected component analysis on regions where the depth values in the physical depth map change abruptly. First, all pixels with depth differences exceeding a set threshold are detected as boundary candidates. Then, these candidate points are connected to form the complete outer contour boundary of the real estate entity. Finally, this outer contour boundary is mapped to the image coordinate system of the corresponding frame in the real estate 3D video through coordinate transformation, and is used as the actual boundary position of the real estate entity for subsequent comparison with the pixel coordinates of the reference boundary position to calculate the average offset distance.
[0050] Specifically, a weighted average method is used to comprehensively determine the modeling accuracy of 3D videos of real estate. In practice, the weight ratio of scale assessment accuracy to boundary assessment accuracy is first determined based on the specific type of real estate and the application scenario requirements of the assessment. Then, the scale assessment accuracy value is multiplied by its weight, and the boundary assessment accuracy value is multiplied by its weight to obtain the comprehensive assessment accuracy value. Finally, this comprehensive assessment accuracy value is compared with a pre-set qualified threshold. When the comprehensive assessment accuracy value reaches or exceeds the qualified threshold, the modeling accuracy is deemed qualified, and a qualified assessment result is output. Otherwise, it is deemed unqualified, and a complete assessment report containing the scale assessment accuracy, boundary assessment accuracy, and comprehensive assessment accuracy values is output.
[0051] It should be noted that the system of this invention adopts a modular architecture design, which is divided into four core modules: ① Input module: receives a single 2D image of real estate and reads the image information; ② Image preprocessing module: adjusts the size, normalizes, and removes noise from the input image to adapt to the input requirements of subsequent models; ③ Estimation module: predicts pixel-level depth information from the 2D real estate image based on the MiDaS model and generates a depth map, providing the core basis for 3D stereoscopic perception; ④ 3D video evaluation module: uses the input image and depth map as conditions to generate 3D-like video frames with camera motion through the SVD model to ensure temporal continuity.
[0052] I. Input Module:
[0053] The input module is the system's data entry point. It receives a single 2D image of real estate provided by the user, performs image format verification, metadata parsing, and basic information extraction, and then passes the standardized data to subsequent modules.
[0054] Specifically, the input data format is defined to receive image files and metadata files. The image files support JPEG and PNG formats, use the RGB color space, and employ three channels with 8 bits of depth per channel. The metadata files are used to transmit the geographic or property rights information of the real estate, and are in JSON or XML format, including: depth scale factor, in meters per pixel; measured dimensions of the real estate, including length, width, and height; property boundary coordinates, including a list of polygon vertex coordinates; and camera intrinsic parameters, including focal length and principal point coordinates.
[0055] Specifically, the image is spatially associated with metadata; that is, image pixels are linked to real-world property boundaries using the property boundary coordinates in the metadata, providing a spatial reference for subsequent modules. The depth scale factor is passed to the depth map refinement module for subsequent geometric correction. The input module encapsulates the processed data into a standard data structure and passes it to the image preprocessing module.
[0056] II. Image Preprocessing Module:
[0057] This module is used to eliminate various distortions and noises in UAV oblique photogrammetry images, enhance architectural detail features, and ensure geospatial consistency. It adapts to the dual input requirements of MiDaS depth estimation and SVD 3D video generation models, and outputs a MiDaS-adapted input image, an SVD-adapted input image, and a preprocessed original reference image. The specific implementation process includes:
[0058] ① Image and geographic information reading: OpenCV is used to read the input oblique photogrammetric image of the real estate, obtain the original resolution and number of channels of the image, and generate a copy of the original image for subsequent comparison and reference; the corresponding POS geospatial data (latitude, longitude, elevation, flight altitude) and camera intrinsic parameters (focal length, principal point coordinates, distortion coefficient) of the image are parsed from the UAV flight log to prepare for subsequent distortion correction.
[0059] ② Oblique photography image distortion correction, which addresses the problems of building outline deformation and spatial scale distortion caused by UAV lens distortion (radial distortion, tangential distortion), and performs distortion correction based on camera intrinsic parameters to ensure the geospatial consistency of real estate images.
[0060] Specifically, before performing distortion correction, two key sets of data need to be extracted: the camera intrinsic parameter matrix and the distortion coefficients. The intrinsic parameter matrix describes the basic projection properties of the camera, and its expression formula is:
[0061] ;
[0062] in, yes The normalized focal length along the axis is derived from the physical focal length. (Unit: mm) Divide by a single pixel Physical width of direction (Unit: mm / pixel) obtained; yes The normalized focal length along the axis is derived from the physical focal length. (Unit: mm) Divide by a single pixel Physical width of direction (Unit: mm / pixel) and It is the key to connecting the real-world scale with the image scale; and The pixel coordinates of the principal point (the intersection of the lens optical axis and the image plane) in the horizontal and vertical directions of the image, usually located near the image center. is the tilt coefficient, representing the degree of tilt of the pixel coordinate axis; ideally, it is 0. In this embodiment, These are the coordinate axes in the camera coordinate system.
[0063] Specifically, the distortion coefficient D is used to describe and correct optical defects in a lens, including radial distortion and tangential distortion, and its expression is: ;in, , , The radial distortion coefficient, caused by the lens shape, results in straight lines at the image edges bending outwards (barrel-shaped) or inwards (pincushion-shaped). The first-order radial distortion coefficient, whose sign directly determines whether it is pincushion distortion ( Image bulging outwards or barrel distortion? The image is concave (on the order of 10). -3 Up to 10 -5 between; These are second-order radial distortion coefficients, primarily affecting the edges and corners of the image, with a magnitude in the range of 10. -5 Up to 10 -7 between; is the third-order radial distortion coefficient, used to correct lenses with extremely large distortion, such as fisheye lenses and ultra-wide-angle lenses. Its value is very small, and it is set to 0 in this embodiment. , The tangential distortion coefficient is... The horizontal tilt of the lens primarily causes the vertical offset. Since the vertical tilt of the lens primarily causes horizontal offset, in UAV mapping, tangential distortion is much smaller than radial distortion. , In the order of 10 -4 Up to 10 -6 between.
[0064] In this embodiment, the two key data extraction methods are based on Zhang Zhengyou's calibration method. Twenty photos of the calibration board (a black and white checkerboard pattern, with each square having a side length of 30mm) are taken from different angles and distances using a drone camera, ensuring the calibration board fills the entire frame and covers all areas of the field of view. For each image, the corner points of the checkerboard are detected, and the corner point coordinates are optimized from pixel-level to sub-pixel-level. Subsequently, this invention constructs an actual coordinate system, setting the checkerboard plane to Z=0, and setting the actual coordinates of the corner points according to their actual dimensions. Finally, the `cv2.calibrateCamera()` function in OpenCV is called, inputting the detected corner point pixel coordinates and their corresponding actual coordinates. The function returns the camera's intrinsic parameter matrix and distortion coefficients.
[0065] Specifically, after obtaining the camera intrinsic parameter matrix and distortion coefficients, the OpenCV cv2.undistort() function is used to perform distortion correction on the image. This function expresses the process of transforming distorted points into distorted points as a mathematical formula based on the camera calibration parameters, and achieves distortion correction by calculating the coordinates of the distorted points. The expression is as follows:
[0066] ;
[0067] in, These are the coordinates of the distorted point in the normalized camera coordinate system (i.e., the uncorrected normalized coordinates). These are the corrected, normalized coordinates. This function takes the original distorted image, camera intrinsic parameter matrix, and distortion coefficient vector as input, and outputs a corrected, distortion-free image, thereby eliminating geometric distortion caused by the camera lens and restoring the true geometric shape of building corners, outlines, and land boundaries.
[0068] ③ Flight strip deviation correction, which addresses the splicing deviation and image offset issues between flight strips in UAV oblique photography. It utilizes the latitude, longitude, and elevation information of POS data to perform slight spatial registration and offset correction on the images within the flight strip.
[0069] Specifically, after distortion correction of a single image is completed, the offset of the current image is calculated using the geographic coordinates of the reference image within the flight strip as a reference. This is a crucial step in achieving accurate image stitching and ensuring geospatial consistency. Simply put, this process utilizes the position information (POS) recorded by the UAV and the camera attitude to accurately project image pixels onto the ground coordinate system. The offset calculation is essentially solving a coordinate transformation problem; both the reference image and the current image have their theoretical positions in the geographic coordinate system, and the difference between the two is the offset.
[0070] Specifically, after the second step of distortion correction, the direction vector in the camera coordinate system of the oblique photographic image is: The rotation from the camera coordinate system to the ground coordinate system is defined by three sub-attitude angles. In aerial photogrammetry, the rotation sequence is yaw first. Then look up and down Finally, roll .
[0071] In this embodiment, the ground coordinate system adopts the ENU coordinate system. The three axes specifically point as follows: E is due east, serving as the X-axis; N is due north, serving as the Y-axis; and U is perpendicular to the horizontal plane, pointing away from the Earth's center, serving as the Z-axis. Thus, the rotation matrix... for: The basic rotation matrix systems are as follows:
[0072] ;
[0073] ;
[0074] ;
[0075] Then the vector in the camera coordinate system After rotation, the vector in the ground coordinate system is obtained. :
[0076] ;
[0077] Specifically, after rotation, this invention needs to determine the intersection point of the light rays and the ground, and the position of the photography center in the ground coordinate system is... Assuming the ground is horizontal and the elevation is constant. The coordinates of the ground point can be obtained as follows:
[0078] ;
[0079] Subsequently, digital elevation models were used. Solve the problem.
[0080] Specifically, take the center point on the image. As a reference point, respectively for the baseline image (subscript) ) and the current image (subscript) Calculate the coordinates of the corresponding ground point:
[0081] ;
[0082] The "pixel to ground" function is implemented through the steps described above (distortion correction, rotation, intersection). Therefore, the offset of the current image relative to the reference image is:
[0083] ;
[0084] To ensure geospatial alignment of multiple images from the same flight path or parcel, an affine transformation is used to correct image translation, preventing stitching misalignment during subsequent batch processing. The affine transformation matrix used in this process is:
[0085] ;
[0086] in, and These are the pixel translations in the horizontal and vertical directions. Applying this matrix to the image pixel coordinates... To obtain the translated coordinates :
[0087] ;
[0088] To align the current image with the reference image, take This moves points in the current image to positions that coincide with corresponding features in the reference image. After obtaining the affine matrix, this invention uses cv2.warpAffine() to resample the current image, generating a translated image. The corrected image retains its original resolution without compromising architectural details or spatial scale.
[0089] ④ Real Estate Building Texture Enhancement: For fine features such as building facades, windows, balconies, and land parcel boundaries in real estate images, texture enhancement is performed before denoising to avoid losing key structural information during subsequent denoising processes, thereby improving the depth estimation's ability to recognize building details.
[0090] Specifically, the Laplacian operator is used to enhance the image's edges, highlighting structural edges such as corners, contours, and doors and windows, allowing the depth estimation model to accurately capture building layers. Let the original color image be... To preserve color information and avoid color cast, this invention converts the image to the YUV color space, enhancing only the luminance component. The conversion formula is as follows:
[0091] ;
[0092] Take the brightness component As the image to be enhanced Since the Laplacian operator is sensitive to noise, this invention first applies a Gaussian filter to the brightness image to suppress high-frequency noise. The Gaussian kernel function used is:
[0093] ;
[0094] The smoothed image is as follows:
[0095] ;
[0096] Then, the Laplace response is calculated using the following formula:
[0097] ;
[0098] The Laplace nucleus is:
[0099] ;
[0100] Output The Laplace response is represented by positive values indicating bright edges on a dark background and negative values indicating dark edges on a bright background. For negative values, this invention uses absolute value normalization to map the Laplace response structure to the range of 0-255 for superposition. The specific process is as follows:
[0101] ;
[0102] ;
[0103] Subsequently, the normalized edge information is superimposed back onto the original brightness image to enhance edge contrast. The enhanced brightness image is as follows:
[0104] ;
[0105] in To enhance the strength coefficient, during the implementation of this invention, it was experimentally measured that when... The effect is best when the pixels are stacked. After stacking, the pixel values need to be truncated to the valid range [0, 255].
[0106] ;
[0107] The enhanced brightness is then merged with the original chroma components and converted back to RGB space.
[0108] ;
[0109] The resulting enhanced color image can be directly input into the depth estimation model.
[0110] Specifically, gamma correction is used for light texture sharpening to enhance texture details in areas such as building walls and park roads, while avoiding over-sharpening that could amplify noise. The gamma correction texture enhancement formula is as follows:
[0111] ;
[0112] in, To enhance the previous pixel value, To enhance the pixel values, The value is the gamma value. For real estate scenarios, it should be between 0.7 and 0.9. The smaller the value, the stronger the sharpening effect.
[0113] ⑤ Noise Removal, Size Adjustment, and Normalization: While preserving architectural details, sensor noise and environmental noise generated during drone photography are eliminated. Gaussian filtering is used to denoise the enhanced RGB images, balancing denoising effectiveness with image detail preservation. Based on subsequent model requirements, the images are adjusted to a specified resolution (384x384 for MiDaS, 1024×576 for SVD). Bicubic interpolation is used to repair missing or distorted pixels, ensuring image clarity and providing a smoother texture mapping foundation for 3D reconstruction. For the different input pixel value requirements of the MiDaS and SVD models, the adjusted images are normalized to eliminate the impact of differences in pixel value magnitude on model inference.
[0114] ⑥ Format Conversion and Saving: Complete the final format conversion of the model input. Images adapted to MiDaS retain the NumPy array format, which is then automatically converted to 4D tensors by the depth estimation module (adding a batch dimension). Images adapted to SVD convert the normalized NumPy arrays to PIL image format to ensure no loss of color and pixel information. Save the preprocessed images for subsequent troubleshooting and performance comparison.
[0115] III. Depth Estimation Module:
[0116] Specifically, considering the characteristics of complex building structures, strong geospatial attributes, and the need to preserve detailed outlines such as building corners, balconies, and bay windows in real estate drone oblique photography images, as well as the need for large-scale multi-strip stitching, improvements were made to the original MiDaS v3.1 (DPT-Hybrid) depth estimation module in five dimensions: model optimization, geospatial constraints, refined post-processing of depth maps, batch adaptation, and accuracy verification. This addresses issues such as the original module's bias in estimating the depth of architectural details in real estate scenes, the lack of geographic scale in depth maps, misalignment in large-scale stitching, and abrupt changes in the depth of fine structures. As a result, the generated depth maps not only meet the stereoscopic requirements of 3D video generation but also conform to the spatial accuracy, structural integrity, and geographic consistency requirements of real estate 3D modeling, providing accurate pixel-level depth data for subsequent real estate 3D modeling accuracy assessment.
[0117] It should be added that, such as Figure 2 The diagram shows the depth estimation model architecture in this embodiment of the invention. The model uses the first processed image as the main input and combines UAV POS geographic data, camera intrinsics, and depth scale factors for inference. The overall network uses the MiDaS v3.1 monocular depth estimation model for images and videos as its base network. The Encoder extracts multi-scale image features, the Decoder recovers pixel-level depth results, and the Head outputs relative depth results. In this invention, the input image undergoes tensor normalization before entering the model. When the resolution exceeds a preset threshold, inference is performed in a block-based manner, and flight strip identification features are added to input units belonging to the same flight strip. During model inference, flight altitude consistency constraints, building structure constraints, edge smoothing constraints, and local second-order structure constraints are further introduced. This ensures that the output depth results not only have visual continuity but also meet the structural expression requirements of real estate scenarios for building outlines, continuous wall areas, corner areas, and land parcel boundary areas, ultimately outputting a visual depth map and physical depth.
[0118] It should be noted that during model training, this invention uses transfer learning to fine-tune the MiDaS model based on oblique photographic images of real estate and corresponding measured depth data of laser point clouds. During training, the Backbone (main encoder) parameters are preferably frozen, and only the Decoder and Head are trained specifically to retain general visual feature extraction capabilities while enhancing their depth representation capabilities for residential, industrial, and commercial real estate scenes. Furthermore, this invention employs a combined loss function to optimize the model. This combined loss includes at least scale-invariant loss, gradient matching loss, and structural similarity loss, and combines flight altitude consistency constraints and building structure constraints on the inference side to achieve result correction. Specifically, scale-invariant loss is used to reduce global scale drift, gradient matching loss is used to enhance the depth abrupt expression of details such as building edges, balconies, and bay windows, and structural similarity loss is used to maintain the overall continuity of local structural regions. This ensures that the resulting depth map provides reliable conditions for subsequent 3D video generation and serves as the physical depth basis for evaluating the accuracy of real estate modeling.
[0119] Specifically, the implementation process of the depth estimation module includes:
[0120] This invention loads a fine-tuned MiDaS v3.1 model based on a real estate oblique photography dataset. The original MiDaS v3.1 model loading logic is modified by adding a fine-tuning weight loading branch to support loading corresponding fine-tuning weights based on real estate type (residential, industrial, and commercial), while retaining general weight loading options. Adding this branch involves inserting a conditional routing module into the existing model loading logic. This module dynamically selects and loads the corresponding fine-tuning weights based on the input real estate type, while retaining general weights as the default option. Since the basic features of real estate images (edges, textures, etc.) are universal, this invention completely freezes the model's backbone encoder, training only the decoder and specific task head for the three real estate types to improve training speed and prevent overfitting. The learning rate for the decoder and head layers is set to 10 times that of the latter half of the backbone network, with an initial learning rate of 0.00001. Furthermore, to optimize fine structures such as building corners, balconies, and bay windows, this invention employs a combined loss function, expressed as:
[0121] ;
[0122] in, The scale-invariant loss function measures the difference between the predicted depth and the true depth in logarithmic space, and removes the influence of global scale shift. Gradient matching loss forces the model to maintain consistency with the true depth at the edges of abrupt depth changes, improving the sharpness of geometric details. As a structural similarity loss, it helps maintain the overall shape and continuity of local areas such as balconies and bay windows, reducing the fragmentation or distortion of the predicted depth map. , , The weights of each loss function, , , After reading the first processed image, tensor normalization is performed on the image. When the resolution of the first processed image exceeds a preset threshold, it is divided into multiple image patches, and each image patch is used as the input unit of the depth estimation model. When the resolution does not exceed the preset threshold, the entire image is used as the input unit for depth estimation inference.
[0123] ② Geospatial parameter reading The process involves reading UAV POS geographic data (latitude, longitude, elevation, flight altitude) and camera intrinsic parameters (focal length, principal point coordinates) to establish a mapping relationship between pixel depth values and actual physical distances (meters), giving the depth map a geographic scale. First, the POS data from the UAV flight log and camera intrinsic parameters are used as input parameters to the depth estimation module and correlated with the preprocessed image. Second, the scale factor is calculated: based on the UAV flight altitude, camera focal length, and image pixel size, the depth scale factor k (unit: meters / pixel) is calculated using the following formula:
[0124] ;
[0125] in, This is the depth scale factor, measured in meters per pixel. For the relative flight altitude of the drone, The physical size of a pixel. This refers to the camera's horizontal focal length.
[0126] ③ Image tensor preprocessing: This involves tensor optimization tailored to the high resolution and multi-strip stitching characteristics of real estate images, ensuring both inference accuracy and speed. First, the tensors of high-resolution real estate images are divided into non-overlapping blocks (512×512) to avoid excessively large single tensors causing memory overflow, while maintaining depth continuity at block junctions. Second, flight strip identifier features (flight strip number encoding) are added to the tensors of real estate images within the same flight strip, allowing the model to maintain spatial consistency during inference. Finally, PrepareForNet() is used for tensor normalization, preserving the 4-dimensional tensor format [batch, channel, H, W] to ensure compatibility with MiDaS model input.
[0127] ④ Constrained Model Inference: This involves introducing flight altitude consistency constraints and building structure constraints during the depth estimation model's inference process. Based on the flight altitude information from the UAV POS geographic data and the focal length information from the camera's intrinsic parameters, the theoretical reference depth for the ground area is determined. Based on the deviation between the theoretical reference depth and the predicted depth of the ground area, flight altitude consistency constraints are constructed to ensure that the predicted depth of the ground area remains consistent with the actual flight altitude. First, the theoretical depth value is calculated based on the flight altitude. Assuming the camera's optical axis is perpendicular to the ground, the ground is a plane parallel to the image plane, at a distance H from the camera. At this point, any pixel... The depth of the corresponding ground point should be:
[0128] ;
[0129] in, This refers to the camera's internal parameters.
[0130] Specifically, let the depth map predicted by the model be... For pixels within the ground region, this invention aims to predict the depth as close as possible to the theoretical depth. Therefore, the flight altitude consistency loss is defined as:
[0131] ;
[0132] in, This represents the total number of ground pixels. For ground mask.
[0133] Specifically, during fine-tuning, the altitude constraint loss is weighted and combined with the original depth estimation loss:
[0134] ;
[0135] in, .
[0136] Secondly, based on the vertical and horizontal structural characteristics of real estate buildings, depth value smoothing constraints are added to avoid abrupt changes in building wall depth and abnormal balcony depth values. This invention defines a depth map. In pixels horizontal gradient at and vertical gradient They are respectively:
[0137] ;
[0138] Specifically, to avoid over-smoothing at real building edges (such as balcony boundaries), this invention introduces a weighting function guided by RGB images. Let The gradient of the luminance channel (or three-channel average) of the original RGB image is:
[0139] ;
[0140] Edge-aware weights are negative exponential functions of the image gradient, which reduces the smoothing penalty for strong edges:
[0141] ;
[0142] in, and Each pixel Edge-sensing weights in both the horizontal and vertical directions; The edge attenuation coefficient is set to 1 in this embodiment to control the edge attenuation rate, resulting in the edge smoothness constraint:
[0143] ;
[0144] in, A first-order directional smoothing constraint for edge sensing; The weights for the horizontal and vertical directions are respectively, both set to 1 in this invention. To further distinguish between planar regions such as wall surfaces and polygonal regions such as wall corners, this invention employs a Laplace constraint with second derivatives. The Laplace operator for the depth map is defined as:
[0145] ;
[0146] Combining edge sensing, the local second-order structure constraint is obtained as follows:
[0147] ;
[0148] in, For depth map at pixels The local second-order structural response at the location; For edge-aware local second-order structure loss; For local second-order structure constraint weights; Edge-aware weights are based on image gradient magnitude. Sure:
[0149] ;
[0150] Finally, the smoothing loss is added as a regularization term to the total loss function:
[0151] ;
[0152] in, These are the first-order directional smoothing constraint weights for edge sensing and the local second-order structural constraint weights for edge sensing, respectively. In this embodiment, we take... After joint optimization based on altitude consistency constraints and building structure constraints, a relative depth result that better conforms to the geometric patterns of real estate scenes can be obtained. This relative depth result maintains the structural consistency of building edges, continuous wall areas, and corner areas, and can be scaled back by combining depth scale factors in subsequent steps, thereby generating a visual depth map for 3D video generation and a physical depth map for modeling accuracy assessment.
[0153] ⑤ Initial post-processing of depth map, that is, adding actual physical depth value conversion on the basis of the original depth map of the model, so that the depth map has both visualization uint8 format and physical depth format.
[0154] Specifically, the depth map tensor is restored to the original image size of the real estate using a bicubic interpolation algorithm. The specific operation is as follows: for the target pixel... Calculate its mapped coordinates in the low-resolution path. Then, a weighted sum is calculated using the depth values of the surrounding 4×4 neighborhood:
[0155] ;
[0156] in, , The weights are the values for the bicubic interpolation kernel. Next, the relative depth values output by the model are scaled using a scaling factor. The formula for converting to actual physical depth (meters) is: ;in, This represents the actual physical depth (meters). The relative depth value (in pixels) output by the model; The depth scale factor is used. Simultaneously, a visual depth map (0-255, uint8, used for 3D video generation) and a physical depth map (meters, float32, used for real estate 3D modeling accuracy assessment) are generated. Finally, the visual depth map is further normalized to [0,255] and converted to uint8 format to ensure compatibility with the input of the 3D video assessment module.
[0157] IV. 3D Video Evaluation Module:
[0158] It should be added that, such as Figure 3The diagram shows the architecture of the 3D video generation model in this embodiment of the invention. The model uses a second processed image and a visualized depth map as basic inputs, and further receives spatial constraint information consisting of depth scale factors, measured dimensions of the real estate, and UAV POS data, as well as property feature parameters consisting of property boundary coordinates, land parcel numbers, and building numbers. Based on this, it combines camera motion parameters to generate a 3D video of the real estate. The model uses a stable video diffusion model (SVD) as the basic generation network, where VAE (Variational Autoencoder) is used for latent space encoding and decoding; U-Net (U-shaped denoising network) is used for noise prediction during the diffusion process; ControlNet (DepthControlNet) is used to inject depth and spatial constraint features into each layer of U-Net; and a spatial attention enhancement module is used to highlight key area features such as building outlines, wall boundaries, and property boundaries.
[0159] It should still be noted that during the training process, such as Figure 3 As shown, this invention adapts the stabilityai / stable-video-diffusion-img2vid-xt pre-trained model for application. By inputting real images and corresponding constraint maps into Depth ControlNet, the model learns the video generation rules under given geometric constraints, thus avoiding architectural scale distortion and structural drift caused by relying solely on two-dimensional textures. The constraint map is constructed from the transformation relationship between image coordinates and actual coordinates, the identification results of wall and ground areas, depth scale factors, and measured dimensions of real estate. During the diffusion sampling process, this invention further introduces scale consistency loss and structural constraint loss to jointly constrain the architectural scale, the vertical / horizontal relationship between walls and ground, and the spatial range of camera pose in the generated frames. For high-resolution large-scene images, a block generation and fusion optimization strategy can also be adopted to reduce memory usage and improve the overall scene consistency.
[0160] Specifically, this invention employs constrained model inference to address the issues of architectural scale distortion and structural deformation that commonly occur in real estate video generation. Specifically, it includes: First, injecting depth scale factors and measured real estate dimensions as constraints into the SVD model generation process to limit the spatial range of camera movement and ensure that the visual proportions of building dimensions in the generated video are consistent with their actual physical dimensions. Second, based on the structural characteristics of the real estate building, adding structural constraints to maintain the vertical and horizontal relationships of the building walls, ground, and roof during model generation. These structural constraints jointly restrict the building wall areas, ground areas, and local depth changes in the generated frames, ensuring that the generated result maintains the vertical and horizontal relationships between the building walls and ground, and suppresses unreasonable deformation in local structural areas.
[0161] Specifically, the total structural constraint loss consists of three components: vertical wall constraint, horizontal plane constraint, and spatial smoothness constraint. The vertical wall constraint penalizes the deviation of the wall normal vector from the horizontal direction, and its expression is:
[0162] ;
[0163] Horizontal plane constraint penalizes deviation of the ground / roof normal vector from the vertical direction:
[0164] ;
[0165] Spatial smoothing constraints employ a depth Laplacian operator to penalize abrupt changes in depth, thus preserving surface smoothness.
[0166] ;
[0167] The Laplace operator It can be discretized and approximated as:
[0168] ;
[0169] In the formula, These are the pixel coordinates in the image coordinate system; For pixels The normal vector of the corresponding surface can be estimated based on the local depth variation relationship of the visualized depth map or the generated frame depth map, and is used to reflect the spatial orientation of the building wall area and the ground area; This is the vertical constraint loss of the wall surface, used to reflect the vertical consistency between the normal vector of the wall surface region and the direction of gravity; This is the horizontal constraint loss, used to reflect the parallelism between the normal vector of the ground or roof area and the direction of gravity; Spatial smoothing constraint loss is used to suppress drastic local fluctuations in the depth map; For the wall area mask, when the pixel If the area belongs to the building wall, use 1; otherwise, use 0. For the ground or rooftop area mask, when the pixel If the area belongs to the ground or rooftop, use 1; otherwise, use 0. The direction vector of gravity; Generate a depth map at pixel points for frame t. The depth value at that location, where t represents the video frame number.
[0170] Specifically, when the resolution of the second processed image exceeds a preset threshold, it is divided into several image sub-blocks as input units for the 3D video generation model. Spatial location identifiers are added to each image sub-block, and corresponding video sub-frame sequences are generated to maintain the positional correspondence of each image sub-block in the overall scene. After brightness and color correction of each video sub-frame sequence, weighted fusion is performed in the overlapping area to eliminate the sub-block splicing boundaries, resulting in a fused video frame sequence. After obtaining the video frame sequence, video synthesis is performed to output a 3D video of real estate.
[0171] Specifically, when evaluating the modeling accuracy of the generated 3D real estate video, key feature points of the real estate are selected based on the physical depth map, combined with camera intrinsics and UAV POS geographic data. The coordinates of each key feature point in 3D space are calculated using backprojection of the physical depth values, and the spatial scale information of the real estate is determined by calculating the Euclidean distance between each key feature point. This spatial scale information is compared with the corresponding measured length, width, and height in the actual measured dimensions of the real estate to calculate the scale evaluation accuracy. Simultaneously, based on the property boundary coordinates, combined with UAV POS geographic data and camera intrinsics, the property boundary coordinates are projected onto the image coordinate system of the corresponding frame in the 3D real estate video to obtain the reference boundary position. The actual boundary position of the real estate is extracted from the 3D real estate video, and the pixel coordinates of the actual boundary position are compared with the reference boundary position to calculate the boundary evaluation accuracy. Finally, the modeling accuracy of the 3D real estate video is comprehensively judged by combining the scale evaluation accuracy and the boundary evaluation accuracy, and the evaluation result is output. Among them, scale assessment accuracy is used to measure the accuracy of the size restoration of the main body of the real estate in the length, width and height directions, and boundary assessment accuracy is used to measure the accuracy of the position fitting of the property boundary in the corresponding frame of the video; the two together reflect the modeling accuracy of the generated real estate 3D video on the actual real estate spatial form and boundary range.
[0172] like Figure 5As shown, according to another embodiment of the present invention, a real estate 3D modeling accuracy evaluation system based on UAV oblique photogrammetry is also provided. This real estate 3D modeling accuracy evaluation system based on UAV oblique photogrammetry includes: an input module, used to extract the image resolution and number of channels of the 2D image based on a 2D image and metadata file of the real estate, and extract the depth scale factor, measured dimensions of the real estate, property boundary coordinates, and camera intrinsic parameters from the metadata file; spatially mapping the property boundary coordinates to the current image resolution; and an image preprocessing module, used to perform oblique photogrammetry correction and enhancement processing on the 2D image according to the 2D image and the corresponding camera intrinsic parameters, and combine this with UAV POS geographic information. The data undergoes flight strip deviation correction to obtain a first processed image adapted to the depth estimation model and a second processed image adapted to the 3D video generation model. The depth estimation module is used to perform depth estimation with geospatial and building structure constraints based on the first processed image, combined with depth scale factors, UAV POS geographic data, and camera intrinsic parameters, to obtain a visual depth map and a physical depth map. The 3D video evaluation module is used to input the second processed image and the visual depth map into the 3D video generation model to generate a 3D video of the real estate, and to evaluate the modeling accuracy of the generated 3D video of the real estate based on the physical depth map, the measured dimensions of the real estate, and the coordinates of the property boundary, and output the evaluation results.
[0173] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for evaluating the accuracy of 3D real estate modeling based on UAV oblique photogrammetry, characterized in that, include: Based on the two-dimensional image and metadata file of the real estate, the image resolution and number of channels of the two-dimensional image are extracted, and the depth scale factor, the measured size of the real estate, the property boundary coordinates and camera intrinsic parameters are extracted from the metadata file; the property boundary coordinates are spatially mapped to the current image resolution; the metadata file is used to convey the geographical or property information of the real estate, including the depth scale factor, the measured size of the real estate, the property boundary coordinates and camera intrinsic parameters; Based on the two-dimensional image and the corresponding camera intrinsic parameters, the two-dimensional image is subjected to oblique photogrammetry correction and enhancement processing, and the flight path deviation is corrected by combining UAV POS geographic data to obtain a first processed image adapted to the depth estimation model and a second processed image adapted to the 3D video generation model. Based on the first processed image, combined with the depth scale factor, UAV POS geographic data and camera intrinsic parameters, depth estimation with geospatial constraints and building structure constraints is performed to obtain a visual depth map and a physical depth map. The second processed image and the visualized depth map are input into the 3D video generation model to generate a 3D video of the real estate. Based on the physical depth map, the measured size of the real estate and the coordinates of the property boundary, the modeling accuracy of the generated 3D video of the real estate is evaluated, and the evaluation result is output. The depth estimation with geospatial and building structure constraints includes: Read the UAV POS geographic data and camera intrinsic parameters corresponding to the first processed image; calculate the depth scale factor based on the flight altitude information in the UAV POS geographic data, the focal length information in the camera intrinsic parameters, and the physical size of the image pixels; perform tensor normalization on the first processed image; when the resolution of the first processed image exceeds a preset threshold, divide the first processed image into multiple image blocks, and use the image blocks as input units of the depth estimation model; when the resolution of the first processed image does not exceed a preset threshold, use the entire first processed image as the input unit of the depth estimation model; add flight strip identification features to input units belonging to the same flight strip, and input the input units into the depth estimation model; introduce flight altitude consistency constraints and building structure constraints during the inference process of the depth estimation model to obtain the constraint-optimized relative depth result; the flight altitude consistency constraints are constructed based on the deviation between the theoretical reference depth and the ground area predicted depth output by the depth estimation model, and the building structure constraints include edge-preserving smoothness constraints and local second-order structure constraints.
2. The method for evaluating the accuracy of 3D real estate modeling based on UAV oblique photography according to claim 1, characterized in that, The process of performing oblique photogrammetry correction and enhancement on the two-dimensional image, and correcting flight strip deviation by combining it with UAV POS geographic data, includes: Based on the two-dimensional image, extract the image resolution and number of channels; The camera intrinsic parameters and UAV POS geographic data corresponding to the two-dimensional image are analyzed; wherein, the camera intrinsic parameters include at least focal length, principal point coordinates, and distortion coefficient; the UAV POS geographic data includes at least latitude and longitude, elevation, and flight altitude; The two-dimensional image is converted into an RGB three-channel image, and the coordinates of the property boundary are spatially mapped to the current image resolution. Using the camera intrinsic parameters and distortion coefficients, distortion correction is performed on the converted RGB three-channel image; based on the UAV POS geographic data, and with the reference image within the flight path as a reference, the geographic coordinate offset of the current image relative to the reference image is calculated; Based on the geographic coordinate offset, the distortion-corrected image is subjected to flight strip deviation correction through affine translation transformation to obtain the processed original reference image; The original reference image is resized, normalized, and converted in format to obtain a first processed image adapted to the depth estimation model and a second processed image adapted to the 3D video generation model.
3. The method for evaluating the accuracy of 3D real estate modeling based on UAV oblique photography according to claim 2, characterized in that, The expression for distortion correction of the converted RGB three-channel image is as follows: ; In the formula, x d and y d The coordinates of the normalized image after correction; x u and y u represents the normalized image coordinates after distortion; k1, k2, and k3 are the radial distortion coefficients; p1 and p2 are the tangential distortion coefficients; The expression for correcting the flight strip deviation is: ; In the formula, u and v are the pixel coordinates of the image after distortion correction; and To complete the image pixel coordinates after flight strip deviation correction; t u and t v These represent the translation amounts of the current image relative to the reference image in the horizontal and vertical directions, respectively.
4. The method for evaluating the accuracy of 3D real estate modeling based on UAV oblique photography according to claim 1, characterized in that, The introduction of flight altitude consistency constraints and architectural structure constraints during the inference process of the depth estimation model, resulting in a constraint-optimized relative depth, includes: The theoretical reference depth of the ground area is determined based on the flight altitude information in the UAV POS geographic data and the focal length information in the camera intrinsic parameters. Construct flight altitude consistency constraints to ensure that the predicted depth of the ground area is consistent with the actual flight altitude; Based on the edge position, continuous wall area, and corner area of the real estate building in the image, the building structure constraint is applied to the relative depth result output by the depth estimation model to suppress depth noise in non-building boundary areas and maintain depth variation at the building outline position. The loss function after introducing the flight altitude consistency constraint and the building structure constraint is jointly optimized, and the relative depth result after constraint optimization is output.
5. The method for evaluating the accuracy of 3D real estate modeling based on UAV oblique photography according to claim 4, characterized in that, The edge-preserving smoothness constraint is based on edge-aware weights determined by the luminance channel gradient of the original RGB image and constructed in combination with the first-order gradients of the depth map in the horizontal and vertical directions, in order to reduce depth fluctuations within the same building plan area. The local second-order structural constraints are constructed based on the local second-order structural response of the depth map and combined with the edge-aware weights, and are used to improve the structural consistency of building walls, balconies and corner areas. The expression for calculating the depth scale factor is as follows: ; In the formula, k is the depth scale factor; H is the relative flight altitude in the UAV POS geographic data; d p f is the physical size of the image pixels; x The horizontal focal length is a camera intrinsic parameter. The expression for the altitude consistency constraint is: ; In the formula, Loss due to flight altitude consistency; This represents the total number of ground pixels. For the ground mask, when the pixel If it belongs to the ground area, use 1; otherwise, use 0. For depth estimation models at pixel points The predicted relative depth at the output; For pixels The theoretical reference depth at that point.
6. The method for evaluating the accuracy of 3D real estate modeling based on UAV oblique photography according to claim 1, characterized in that, The process of generating 3D videos of real estate includes: Read the second processed image and the visualized depth map; Obtain the depth scale factor, actual size of real estate, UAV POS geographic data and property rights feature parameters corresponding to the second processed image, and construct spatial constraint information based on the UAV POS geographic data and actual size of real estate; The camera motion parameters are set according to the preset real estate camera motion mode, and the second processed image, the visualization depth map, the spatial constraint information and the camera motion parameters are input into the 3D video generation model to perform constrained model inference to obtain a video frame sequence. When the resolution of the second processed image exceeds a preset threshold, the video frame sequence is divided into blocks and fused. The generated video frame sequence is processed for video synthesis and optimization, and a 3D video of the real estate is output.
7. The method for evaluating the accuracy of 3D real estate modeling based on UAV oblique photography according to claim 6, characterized in that, The process of performing constrained model inference to obtain the video frame sequence includes: Based on the depth scale factor and the measured size of the real estate, establish the scale correspondence between the generated frame depth map and the actual physical size of the real estate; The scale correspondence is introduced as a scale constraint into the diffusion sampling process of the 3D video generation model to limit the range of scale changes of the building body during the generation process. The building wall area and ground area are determined based on the visualized depth map, and structural constraints are constructed based on the normal relationship between the building wall area and ground area; The structural constraints are incorporated into the diffusion sampling process to maintain the vertical and horizontal structural relationship between the building walls and the ground. Based on the measured dimensions of the real estate and the scene center point, spatial range constraints are applied to the camera pose corresponding to each frame during the diffusion sampling process. Based on the diffusion sampling results after introducing the scale constraint, the structural constraint and the spatial range constraint, the video frame sequence is output.
8. The method for evaluating the accuracy of 3D real estate modeling based on UAV oblique photography according to claim 6, characterized in that, The accuracy assessment of the generated 3D real estate video modeling includes: Based on the physical depth map, combined with camera intrinsic parameters and UAV POS geographic data, key feature points of the real estate are selected. The coordinates of each key feature point in three-dimensional space are calculated by back-projection of the physical depth values. The spatial scale information of the real estate is determined by calculating the Euclidean distance between the coordinates of each key feature point in three-dimensional space. The spatial scale information includes the length, width and height of the real estate. The spatial scale information is compared with the actual measured length, width and height of the real estate, and the relative error percentage of each dimension is calculated. The average of the three is taken as the scale assessment accuracy. Based on the property boundary coordinates, combined with UAV POS geographic data and camera intrinsic parameters, the property boundary coordinates are projected onto the image coordinate system of the corresponding frame of the real estate 3D video to obtain the reference boundary position. Extract the actual boundary position of the real estate entity from the 3D video of the real estate, compare the pixel coordinates of the actual boundary position with the reference boundary position, and calculate the average offset distance as the boundary assessment accuracy. By combining the scale assessment accuracy and the boundary assessment accuracy, the modeling accuracy of the real estate 3D video is comprehensively judged, and the corresponding assessment result is output.
9. A real estate 3D modeling accuracy evaluation system based on UAV oblique photogrammetry, used to implement the real estate 3D modeling accuracy evaluation method based on UAV oblique photogrammetry as described in any one of claims 1-8, characterized in that, The system includes: The input module is used to extract the image resolution and number of channels of the two-dimensional image based on the real estate's two-dimensional image and metadata file, and to extract the depth scale factor, measured size of the real estate, property boundary coordinates, and camera intrinsic parameters from the metadata file; and to spatially map the property boundary coordinates to the current image resolution; the metadata file is used to transmit the geographic or property information of the real estate, including the depth scale factor, measured size of the real estate, property boundary coordinates, and camera intrinsic parameters; The image preprocessing module is used to perform oblique photogrammetry correction and enhancement processing on the two-dimensional image based on the two-dimensional image and the corresponding camera intrinsic parameters, and to perform flight strip deviation correction in combination with UAV POS geographic data to obtain a first processed image adapted to the depth estimation model and a second processed image adapted to the 3D video generation model. The depth estimation module is used to perform depth estimation with geospatial constraints and building structure constraints based on the first processed image, combined with the depth scale factor, UAV POS geographic data and camera intrinsic parameters, to obtain a visual depth map and a physical depth map. The 3D video evaluation module is used to input the second processed image and the visualized depth map into the 3D video generation model to generate a 3D video of the real estate, and to evaluate the modeling accuracy of the generated 3D video of the real estate based on the physical depth map, the measured size of the real estate and the coordinates of the property boundary, and output the evaluation results. The depth estimation with geospatial and architectural constraints includes: reading UAV POS geographic data and camera intrinsic parameters corresponding to the first processed image; calculating a depth scale factor based on the flight altitude information in the UAV POS geographic data, the focal length information in the camera intrinsic parameters, and the physical size of the image pixels; performing tensor normalization on the first processed image; dividing the first processed image into multiple image blocks when the resolution exceeds a preset threshold, and using the image blocks as input units of the depth estimation model; using the entire first processed image as the input unit of the depth estimation model when the resolution does not exceed the preset threshold; adding flight strip identification features to input units belonging to the same flight strip, and inputting the input units into the depth estimation model; introducing flight altitude consistency constraints and architectural constraints during the inference process of the depth estimation model to obtain a constraint-optimized relative depth result; the flight altitude consistency constraints are constructed based on the deviation between the theoretical reference depth and the predicted depth of the ground area output by the depth estimation model, and the architectural constraints include edge-preserving smoothness constraints and local second-order structure constraints.
Citation Information
Patent Citations
Unmanned aerial vehicle scene dense reconstruction method based on VI-SLAM and depth estimation network
CN112435325A
Electronic anti-shake method and device based on single shooting depth perception and storage medium
CN122093664A