A method and system for three-dimensional reconstruction of urban buildings based on virtual modeling

The multi-view image of urban buildings is obtained through drones, the photometric compensation weight is calculated and the deep learning model is trained, which solves the problem of inactivity of the photometric consistency assumption in traditional multi-view three-dimensional reconstruction and improves the accuracy of the three-dimensional reconstruction model.

CN119579805BActive Publication Date: 2025-06-06XINZHIHANG MEDIA TECH GRP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510138700.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-06-06
Estimated Expiration
2045-02-08

AI Technical Summary

Technical Problem

The traditional multi-view three-dimensional reconstruction technology is based on the photometric consistency assumption, but in reality, it is affected by light conditions and building surface characteristics, resulting in the photometric consistency assumption that is invalid, causing noise to increase, affecting the accuracy of the three-dimensional reconstruction model.

Method used

The drone obtains the grayscale map and depth map from different perspectives of urban buildings, calculates the depth homogeneity and reflection standard differences of each pixel point, combines the grayscale deviation to calculate the photometric compensation weight, and builds a loss function to train a deep learning model to obtain a three-dimensional model of urban buildings.

Benefits of technology

The accuracy of the depth map and the accuracy of the three-dimensional reconstruction model are improved, and the deviation problem of the photometric consistency assumption in the actual environment is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119579805B_ABST
    Figure CN119579805B_ABST
Patent Text Reader

Abstract

The present application relates to the field of image enhancement three-dimensional reconstruction technology, and specifically to a method and system for three-dimensional reconstruction of urban buildings based on virtual modeling, the method comprising: using a camera mounted on an unmanned aerial vehicle to obtain grayscale images and depth images of urban buildings at different viewing angles; obtaining the depth homogeneity of each pixel in each depth image based on the similarity of the depth value changes in the local window of each pixel in each depth image; obtaining the photometric compensation weight of each pixel in the grayscale image of each viewing angle based on the spatial difference and grayscale distribution difference of point cloud data between each viewing angle and its adjacent viewing angle, and the grayscale difference of the pixels at the same position in the grayscale image of each viewing angle and its adjacent viewing angle, in combination with the reflection standard difference and the depth homogeneity, so as to construct a loss function, train a deep learning model and obtain a three-dimensional model of the urban building. The present application improves the accuracy of three-dimensional reconstruction of virtual modeling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image enhancement three-dimensional reconstruction, and in particular to a method and system for three-dimensional reconstruction of urban buildings based on virtual modeling. Background Art

[0002] The three-dimensional reconstruction of urban buildings based on virtual modeling refers to the technology of using computer technology and plane geometry technology to construct real urban buildings in a three-dimensional digital form in a virtual environment. It helps to directly present the building structure, visualize the exhibition hall representation, and improve the safety of building construction. It has high potential in urban planning and design, building development and protection, etc. In the three-dimensional reconstruction of urban buildings, laser scanners or structured light equipment are used to actively scan the buildings to achieve three-dimensional reconstruction of urban buildings. However, active three-dimensional reconstruction has high requirements for equipment hardware and high cost of use, which is not conducive to its popularization and use. Therefore, the three-dimensional reconstruction technology based on passive multi-view has higher potential. It only needs to obtain multi-angle images of urban buildings through ordinary cameras and reconstruct the three-dimensional model of the object using computer technology.

[0003] The 3D reconstruction of images is usually based on the photometric consistency assumption, that is, the color texture of the same point on the real building surface is consistent under different camera viewing angles. However, in reality, due to the influence of lighting conditions and the reflection differences of the building surface characteristics, the color of the same point under different viewing angles may be different, which cannot meet the photometric consistency assumption, resulting in more noise in the point cloud data of the 3D reconstruction of urban buildings, affecting the accuracy of the virtual model of urban buildings. Summary of the invention

[0004] In order to solve the above technical problems, the purpose of this application is to provide a method and system for three-dimensional reconstruction of urban buildings based on virtual modeling. The technical solutions adopted are as follows:

[0005] The embodiment of the present application provides a method for three-dimensional reconstruction of urban buildings based on virtual modeling, comprising the following steps:

[0006] Use the camera mounted on the drone to obtain grayscale images and depth images of urban buildings from different perspectives;

[0007] Obtain a local window of each pixel, and obtain the depth homogeneity of each pixel in each depth map based on the similarity of the depth value changes within the local window of each pixel in each depth map;

[0008] The depth map of each viewing angle is converted to obtain the point cloud data corresponding to each viewing angle. Based on the spatial difference and grayscale distribution difference of each viewing angle with respect to the point cloud data, the reflection standard difference of each pixel in the depth map under each viewing angle is constructed. According to the grayscale difference of the pixel at the same position in the grayscale image of each viewing angle and its adjacent viewing angle, the grayscale deviation of each pixel in the grayscale image under each viewing angle is constructed. Combined with the reflection standard difference and the depth homogeneity, the photometric compensation weight of each pixel in the grayscale image of each viewing angle is obtained.

[0009] A loss function is constructed based on the photometric compensation weights to train the deep learning model and obtain a three-dimensional model of urban buildings.

[0010] Preferably, the grayscale image and the depth image at the same viewing angle have the same size, and the pixels correspond one to one.

[0011] Preferably, the local window is a square window with each pixel as the center and a preset side length.

[0012] Preferably, the calculation formula for the depth homogeneity is:

[0013] ; In the formula, Indicates the depth homogeneity of pixel X in the depth map of view 1, N represents the number of pixels in the local window, and cos() represents the calculation of the cosine similarity coefficient. They respectively represent the gradient vectors of the ith pixel and the jth pixel in the local window of the pixel X, and the gradient vector of each pixel is obtained by transforming the gradient of the depth map.

[0014] Preferably, the calculation formula corresponding to the reflection standard difference is:

[0015] ; In the formula, is the reflection standard difference of pixel X in the depth map of view 1, is the radian value of the angle between the planes corresponding to view 1 and view 2, To avoid constants with zero denominators, is the point cloud pair corresponding to view 1 and view 2 ( ), wherein the point cloud matching algorithm matches the point cloud data corresponding to perspective 1 and perspective 2 to obtain a point cloud pair, perspective 2 and perspective 1 are two adjacent perspectives, and the pixel point corresponding to the point cloud x in the point cloud data of perspective 1 in the depth map is pixel point X.

[0016] Preferably, the optical center of view 1 and the optical center of view 2 are used to construct a plane with the point cloud x under view 1, and accordingly, the optical center of view 1 and the optical center of view 2 are used to construct a plane with the point cloud x under view 2. Construct a plane where point cloud x and point cloud is a pair of matched point clouds.

[0017] Preferably, the grayscale deviation corresponds to the calculation formula:

[0018] ; In the formula, Indicates the grayscale deviation of pixel Y in the grayscale image of view 1, Indicates the number of pixels in the local window, based on the actual number of pixels in the local window. Represents pixel pairs The absolute value of the grayscale difference between two pixels. Represents pixel pairs The absolute value of the grayscale difference between the pixels at the t-th position in the two local windows;

[0019] Among them, the perspective 1 corresponds to the depth map a and the grayscale map a, and the perspective 2 corresponds to the depth map b and the grayscale map b. The two pixel points of the point cloud pairs corresponding to the perspective 1 and the perspective 2 in the depth map a and the depth map b are recorded as depth point pairs, and the two pixel points corresponding to the depth point pairs in the grayscale map a and the grayscale map b constitute pixel pairs.

[0020] Preferably, the corresponding calculation formula of the luminosity compensation weight is:

[0021] ; In the formula, represents the photometric compensation weight of pixel k in the grayscale image of view 1, norm() represents the normalization function, represents the reflection standard difference of pixel k in the depth map of view 1, represents the depth homogeneity of pixel k in the depth map of view 1, Represents the grayscale deviation corresponding to pixel k in the grayscale image of view 1.

[0022] Preferably, the constructing of the loss function based on the photometric compensation weight further comprises:

[0023] The difference between the depth map obtained by the deep learning model at each viewing angle and the depth value of the corresponding position in the label depth map is calculated. The differences obtained at all positions constitute the depth difference matrix at each viewing angle. The L1 norm of the product between the photometric compensation weight of each pixel point in the grayscale image of each viewing angle and the depth difference matrix of the corresponding viewing angle is used as the loss function.

[0024] An embodiment of the present application also provides a three-dimensional reconstruction system of urban buildings based on virtual modeling, including a memory, a processor, and a computer program stored in the memory and running on the processor, and when the processor executes the computer program, the steps of any one of the above methods are implemented.

[0025] From the above, it can be seen that the method and system for 3D reconstruction of urban buildings based on virtual modeling provided by the present application have at least the following beneficial effects:

[0026] This application obtains a dataset for 3D reconstruction of urban buildings through a drone acquisition platform. In the multi-view 3D reconstruction training process based on deep learning, the depth homogeneity of a single pixel is obtained based on the characteristics of the depth distribution of pixels in the depth map under a single view, reflecting the degree to which a single pixel is affected by the luminosity. At the same time, based on the spatial difference of a single point cloud pair under adjacent views, the difference in reflection representation is obtained, reflecting the accuracy of the depth map of the point cloud pair under the influence of building surface features. Based on this, the grayscale deviation is obtained based on the grayscale distribution difference under adjacent views, reflecting the confidence of the building surface features in the range of the luminosity consistency area under different viewing angles. On this basis, the luminosity compensation weight of a single pixel is obtained, and the luminosity compensation weight is obtained by traversing each pixel under a single view, and a loss function is constructed to realize model training.

[0027] This application solves the problem that traditional multi-view 3D reconstruction is based on the assumption of photometric consistency. However, in practice, the assumption of photometric consistency is biased by the influence of light and building surface reflection characteristics, which ultimately affects the accuracy of the depth map. This application focuses on analyzing the characteristics of photometric in depth map and grayscale in a single view and in adjacent views, obtaining photometric compensation weights, correcting the original loss function, and improving the accuracy of the depth map and the accuracy of the reconstructed model. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0029] Figure 1 A flowchart of the steps of a method for three-dimensional reconstruction of urban buildings based on virtual modeling provided in this application. DETAILED DESCRIPTION

[0030] In order to further explain the technical means and effects adopted by the present application to achieve the predetermined invention purpose, the following, in combination with the accompanying drawings and preferred embodiments, describes in detail a method and system for three-dimensional reconstruction of urban buildings based on virtual modeling proposed in the present application, its specific implementation method, structure, features and effects. In the following description, different "one embodiment" or "another embodiment" does not necessarily refer to the same embodiment. In addition, specific features, structures or characteristics in one or more embodiments may be combined in any suitable form.

[0031] Unless otherwise specified and limited, terms such as "comprises", "comprising" or any other variants thereof are intended to cover non-exclusive inclusion, so that a circuit structure, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such articles or devices. In the absence of further restrictions, an element defined by the sentence "comprising a ..." does not exclude the presence of other identical elements in the article or device including the element. In addition, the term "and\or" used herein includes any and all combinations of one or more related listed items. All technical and scientific terms used herein have the same meaning as those commonly understood by technicians in the technical field of this application.

[0032] The following is a detailed description of a method and system for three-dimensional reconstruction of urban buildings based on virtual modeling provided by the present application in conjunction with the accompanying drawings.

[0033] See also Figure 1 , which shows a flowchart of a method for three-dimensional reconstruction of urban buildings based on virtual modeling provided by an embodiment of the present application, including the following steps:

[0034] Step 1: Use the camera mounted on the drone to obtain grayscale images and depth images of urban buildings from different perspectives.

[0035] In traditional multi-view 3D reconstruction, the corresponding depth map is constructed based on the derivation of plane geometry through image data from each viewing angle, and the point cloud data of the scene is constructed in combination with camera parameters, that is, a 3D model is obtained. Although the traditional multi-view 3D reconstruction effect is good, the requirements for model texture are high in the actual process. Therefore, this embodiment mainly focuses on multi-view 3D reconstruction based on deep learning.

[0036] In order to obtain model data for deep learning, this embodiment uses a drone mounting platform equipped with a high-definition industrial camera to obtain image data of a single building in the city from different perspectives. In order to obtain sufficient data, this embodiment obtains 100 pictures of a single city building, and a total of 500 city buildings are collected, and the corresponding camera parameter information is obtained, including the camera's intrinsic and extrinsic parameters.

[0037] Since the deep learning model is a supervised training, it is necessary to obtain labeled data. Therefore, for each building at each viewing angle, a laser scanner is used to obtain the depth map of the corresponding viewing angle as the label depth map to obtain the label data. It should be noted that the labeled data only participates in the calculation of the loss function during the model training process. After the deep learning model test and the final training of the deep learning model are completed, it is only necessary to input the multi-view images of the urban building into the trained deep learning model to obtain the corresponding three-dimensional model of the urban building, that is, the point cloud data.

[0038] Step 2: Obtain a local window of each pixel, and obtain the depth homogeneity of each pixel in each depth map based on the similarity of the depth value changes within the local window of each pixel in each depth map.

[0039] Based on the above process, a data set based on urban buildings can be obtained. In this embodiment, 80% of the data set is used as a training set, and the remaining 20% ​​is used as a test set, thereby realizing the division of the data set into a training set and a test set.

[0040] In this embodiment, the three-dimensional reconstruction of urban buildings is mainly based on the end-to-end deep learning model MVSNet, which mainly includes a multi-scale feature extraction module, a region matching module, a cost aggregation module, and a depth regression module. However, when constructing the loss function in the model training and learning process, only the L1 loss function corresponding to the depth map and the annotated depth map under each viewing angle is output, and the weights of each neuron in the convolutional model are updated by back propagation.

[0041] However, this model is still based on the assumption of photometric consistency. However, in practice, it is affected by ambient light and the reflective characteristics of urban buildings, which will cause the depth map of the model in the actual environment to be quite different from the actual depth map, affecting the accuracy of the final three-dimensional model of urban buildings. It is necessary to compensate for photometric consistency during the training process in combination with the training set.

[0042] The following analysis is based on image data of a single urban building scene building. In the three-dimensional reconstruction process, the image is converted from a two-dimensional pixel coordinate system to a three-dimensional camera coordinate system using the camera's intrinsic parameter matrix, mainly based on a single perspective. The camera coordinate system is a three-dimensional space constructed with the camera optical center of the current perspective camera as the coordinate origin, and is converted to a world coordinate system using rotation and translation through the camera's extrinsic parameter matrix. The world coordinate system is also a three-dimensional coordinate system obtained by rotating and translating the camera coordinate system.

[0043] Based on a single viewing angle, the MVSNet model is used to finally output a depth map of the same size under that viewing angle based on the depth regression module. The value corresponding to each pixel in the depth map represents the distance from the real point on the building surface to the optical center of the camera under the current viewing angle.

[0044] When the photometric consistency assumption does not hold, the depth map corresponding to the pixels in the area affected by the lighting may be biased. Specifically, in urban buildings, the geometric structure of the surface is often relatively regular, basically a combination of flat surfaces. Therefore, in the ideal depth map, the gradients of the depth values ​​corresponding to each flat surface are basically the same. If the depth gradient of a certain pixel point is significantly different from the surrounding neighboring pixels, it indicates that the depth value corresponding to the pixel point is likely to be interfered by light, indicating that the accuracy of the depth value at this position is poor.

[0045] Therefore, in this embodiment, the depth map corresponding to each viewing angle is subjected to gradient transformation to obtain the gradient vector of each pixel. It should be noted that the process of gradient calculation of the depth map is an existing well-known technology and will not be described in detail in this embodiment. Furthermore, in this embodiment, a local window is constructed with a single pixel as the center. Based on the above analysis, the depth homogeneity of each pixel in the depth map of different viewing angles is constructed. Taking any viewing angle as viewing angle 1 as an example, the specific calculation formula in this embodiment is:

[0046] ; In the formula, Indicates the depth homogeneity of pixel X in the depth map of view 1, N represents the number of pixels in the local window, and for pixels close to the boundary, the actual number of pixels in the local window shall prevail. cos() represents the calculation of the cosine similarity coefficient. Respectively represent the gradient vectors of the i-th pixel point and the j-th pixel point in the local window of the pixel point X.

[0047] It is understandable that if the current pixel in the depth map is greatly affected by the luminosity, the depth value difference in the local neighborhood of the pixel fluctuates greatly, the gradient difference of the depth value in the local neighborhood is large, the cosine similarity is small, and the depth homogeneity value of the pixel is finally small. On the contrary, if the current pixel is less affected by the luminosity, the depth gradient in the local neighborhood is basically consistent, and the depth homogeneity value is large.

[0048] Step 3: Convert the depth map of each viewing angle to obtain the point cloud data corresponding to each viewing angle. Based on the spatial difference and grayscale distribution difference of the point cloud data between each viewing angle and its adjacent viewing angle, construct the reflection standard difference of each pixel in the depth map at each viewing angle. According to the grayscale difference of the pixels at the same position in the grayscale image of each viewing angle and its adjacent viewing angle, construct the grayscale deviation of each pixel in the grayscale image at each viewing angle. Combine the reflection standard difference and the depth homogeneity to obtain the photometric compensation weight of each pixel in the grayscale image of each viewing angle.

[0049] Based on the distribution difference of the depth map gradient corresponding to each viewing angle, the depth homogeneity of a single pixel is obtained, which reflects the degree of luminosity influence of the local range of the pixel in the depth map. During the camera shooting process, affected by the reflection characteristics of the object surface, it is more likely to be distributed in a regional range, affecting the accuracy of the luminosity consistency compensation of the depth map under a single viewing angle.

[0050] Regarding the image data at each viewing angle, this embodiment takes viewing angle 1 as an example for detailed explanation, and performs a joint analysis based on the depth and grayscale features of the pixel points in the image data of its adjacent viewing angles to evaluate the luminous influence of the pixel points at each position in the image data at each viewing angle.

[0051] In this embodiment, the adjacent perspective of perspective 1 is recorded as perspective 2, and perspective 1 and perspective 2 are two adjacent perspectives. It should be noted that this embodiment performs annotation order during the image data acquisition process at different perspectives, and the implementer can annotate by himself in the actual scene. In this embodiment, for the convenience of analysis, the depth maps of the two adjacent perspectives of perspective 1 and perspective 2 are represented by depth map a and depth map b respectively. The corresponding depth maps under the two perspectives are different, but they correspond to the same building. Therefore, the motion structure recovery SFM method is used to convert the pixels in the depth map into point cloud data in the world coordinate system in combination with the camera's external parameter matrix. It should be noted that the process of obtaining point cloud data through the depth map is an existing well-known technology, which will not be repeated in this embodiment. In actual application scenarios, the implementer can choose other existing methods for conversion.

[0052] Since the depth maps under different viewing angles may have deviations, the point clouds converted under different viewing angles also have differences. Therefore, the point cloud data under viewing angle 1 is paired with the point cloud data under viewing angle 2 to obtain a point cloud pair, such as , thus, mapping to the corresponding pixel pairs in depth map a and depth map b , recorded as a depth point pair. It should be noted that there are many point cloud matching methods, such as ICP (Iterative Closest Point) and NDT (Normal Distributions Transform) algorithms. In this embodiment, the ICP point cloud matching algorithm is used, and the implementer can choose it in the actual application scenario. The specific process of point cloud matching is an existing well-known technology and will not be repeated in this embodiment.

[0053] Furthermore, in two adjacent viewing angles, plane 1 is constructed based on the optical center of viewing angle 1, the optical center of viewing angle 2, and the point cloud x under viewing angle 1. Similarly, plane 1 is constructed based on the optical center of viewing angle 1, the optical center of viewing angle 2, and the point cloud x under viewing angle 2. , construct plane 2, where point cloud x and point cloud The smaller the angle between the two planes, the higher the accuracy of the pairing of the two point clouds. If the distance between the point cloud pairs in the world coordinate system is greater, it means that the depth point pairs in the depth map are more affected by the reflection characteristics of the object surface.

[0054] Based on the above analysis, the reflection representation difference of each pixel point in the depth map under each viewing angle is constructed. View 1 and view 2 are two adjacent viewing angles. In this embodiment, the specific calculation formula is:

[0055] ; In the formula, Represents the reflection standard difference of pixel X in the depth map of view 1, Represents the point cloud pair corresponding to view 1 and view 2 The Euclidean distance between two point clouds in represents the angle between the planes corresponding to view 1 and view 2, where the unit of the angle is radian. To avoid a constant with a denominator of zero, the value range is 0 to 0.1. =0.001.

[0056] Ideally, the dual planes constructed under adjacent viewing angles overlap, indicating that the optical center of adjacent viewing angles and the point cloud are in the same plane, which is called the epipolar plane in epipolar geometry. In practice, if the difference between the two is smaller, it indicates that the matching point clouds of the two represent the same position in the actual building. The greater the spatial difference between the two corresponding point cloud pairs, the greater the influence of the surface characteristics of the object on the depth map under the two viewing angles. The greater the value of the reflection representation difference of the pixel point in the depth map corresponding to viewing angle 1.

[0057] In addition, under the influence of the surface characteristics of the object, the actual grayscale image mainly presents a regional distribution. That is, at different viewing angles, if affected by light and the surface of the object, although the color of the same point of the building is different at different viewing angles, the color difference is consistent in the local range.

[0058] Based on the above analysis, each depth point pair is mapped to the grayscale image, where the position of each pixel in the depth image corresponds to the position of each pixel in the grayscale image, and the depth point pairs under depth image a and depth image b are mapped to the grayscale image. The corresponding pixel pairs in grayscale image a and grayscale image b are recorded as pixel pairs , where grayscale image a is a grayscale image at the same viewing angle as depth image a, grayscale image b is a grayscale image at the same viewing angle as depth image b, and viewing angles 1 and 2 are two adjacent viewing angles. Similarly, a local window of each pixel in the grayscale image is constructed, and a grayscale deviation is constructed based on the difference in grayscale changes within the local window between the corresponding matching pixel pairs in the grayscale image at each viewing angle and its adjacent viewing angle. In this embodiment, the corresponding calculation formula is:

[0059] ; In the formula, Indicates the grayscale deviation of pixel Y in the grayscale image of view 1, Indicates the number of pixels in the local window, based on the actual number of pixels in the local window. Represents pixel pairs The absolute value of the grayscale difference between two pixels. Represents pixel pairs The absolute value of the grayscale difference between the pixels at the t-th position in the two local windows.

[0060] When the surface features of an object are affected at each viewing angle, the impact often presents a regional distribution, and the regional changes are consistent. Therefore, under the local window, the smaller the grayscale difference between the pixel point in the local window and the central pixel point, the smaller the grayscale deviation value is. Therefore, based on the reflection standard difference and depth homogeneity of the pixel points at each position, and combined with the grayscale deviation, the photometric compensation weight of each pixel point in the grayscale image of each viewing angle is constructed. Similarly, taking viewing angle 1 as an example, the specific calculation formula in this embodiment is:

[0061] ; In the formula, represents the photometric compensation weight of pixel k in the grayscale image of view 1, norm() represents the normalization function, represents the reflection standard difference of pixel k in the depth map of view 1, represents the depth homogeneity of pixel k in the depth map of view 1, It represents the grayscale deviation corresponding to the pixel k in the grayscale image of view 1. It should be noted that the depth map and the grayscale image have the same size and the pixel positions correspond one to one under the same view angle.

[0062] It can be understood that when the influence of photometric data is greater, the spatial distance between the corresponding point clouds of pixels under adjacent viewing angles is greater, that is, the difference in reflection standards is greater. At the same time, the greater the gradient difference in the depth map, the smaller the corresponding depth homogeneity, the greater the grayscale deviation in the grayscale map, and the greater the final photometric compensation weight, that is, the greater the loss value of the pixel when calculating the loss function. Therefore, in order to ensure that the pixel with greater influence on photometric data during model training has a greater learning tendency.

[0063] Step 4: Construct a loss function based on the photometric compensation weights to train the deep learning model and obtain a three-dimensional model of urban buildings.

[0064] Based on the above process, the photometric compensation weight of each pixel in the grayscale image under each viewing angle can be obtained, thereby traversing each pixel in the grayscale image under each viewing angle to obtain the photometric compensation weight matrix corresponding to each viewing angle. When calculating the loss function during the training process of the deep learning model MVSNet, in this embodiment, the difference between the depth map reconstructed under a single viewing angle and the depth value of the corresponding position in the label depth map under the viewing angle is calculated to obtain the depth difference matrix under a single viewing angle. Therefore, the L1 norm of the product of the photometric compensation weight of each pixel in the grayscale image of each viewing angle and the depth difference matrix of the corresponding viewing angle is used as the loss function to train the deep learning model MVSNet. It should be noted that for a single building, when analyzing the last viewing angle, the first viewing angle is used as the adjacent viewing angle of the last viewing angle. In this embodiment, the optimizer during the training of the deep learning model MVSNet selects the Adam optimizer for iterative training. The selection of the optimizer is selected by the implementer in the actual application scenario. The specific process of model training is an existing well-known technology and will not be repeated in this embodiment.

[0065] Thus, the model is trained using the training set, and the model is tested using the test set. After the deep learning model training is completed, the image data of the collected building can be used to output the depth map of the corresponding building according to the trained model in this embodiment. The depth map can be converted into a corresponding point cloud data model, that is, a three-dimensional model of the building.

[0066] Based on the same inventive concept as the above method, an embodiment of the present application also provides a three-dimensional reconstruction system of urban buildings based on virtual modeling, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of any one of the above-mentioned methods for three-dimensional reconstruction of urban buildings based on virtual modeling are implemented.

[0067] It is to be understood that the sequence of the embodiments of the present application described above is for description only and does not represent the advantages and disadvantages of the embodiments. The above describes specific embodiments of the present specification. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0068] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. The above content is only an implementation method of this application and is not used to limit the scope of this application. Any equivalent structure or equivalent process transformation made using the content of this application specification and drawings, or directly or indirectly used in other related technical fields, is also included in the protection scope of this application.

Claims

1. A method for three-dimensional reconstruction of urban buildings based on virtual modeling, characterized in that: The following steps are involved: Use the camera mounted on the drone to obtain grayscale images and depth images of urban buildings from different perspectives; Obtain a local window of each pixel, and obtain the depth homogeneity of each pixel in each depth map based on the similarity of the depth value changes within the local window of each pixel in each depth map; The depth map of each viewing angle is converted to obtain the point cloud data corresponding to each viewing angle, and the reflection standard difference of each pixel point in the depth map under each viewing angle is constructed. The calculation formula is: ; In the formula, is the reflection standard difference of pixel X in the depth map of view 1, is the radian value of the angle between the planes corresponding to view 1 and view 2, To avoid constants with zero denominators, is the point cloud pair corresponding to view 1 and view 2 The Euclidean distance between two point clouds in the image is obtained by matching the point cloud data corresponding to the view 1 and the view 2 by the point cloud matching algorithm, wherein the point cloud pair is obtained by matching the point cloud data corresponding to the view 1 and the view 2, the view 2 and the view 1 are two adjacent view angles, and the pixel point corresponding to the point cloud x in the view 1 point cloud data in the depth map is the pixel point X; according to the grayscale difference of the pixel points at the same position in the grayscale image of each view angle and its adjacent view angle, the grayscale deviation of each pixel point in the grayscale image under each view angle is constructed, and the photometric compensation weight of each pixel point in the grayscale image of each view angle is obtained by combining the reflection standard difference and the depth homogeneity; A loss function is constructed based on the photometric compensation weights to train the deep learning model and obtain a three-dimensional model of urban buildings.

2. A method for three-dimensional reconstruction of urban buildings based on virtual modeling as claimed in claim 1, characterized in that: The grayscale image and depth image at the same viewing angle have the same size and the pixels correspond one to one.

3. The method for three-dimensional reconstruction of urban buildings based on virtual modeling according to claim 1, characterized in that: The local window is a square window with each pixel as the center and a preset side length.

4. The method for three-dimensional reconstruction of urban buildings based on virtual modeling according to claim 1, characterized in that: The calculation formula of the depth homogeneity is: ; In the formula, Indicates the depth homogeneity of pixel X in the depth map of view 1, N represents the number of pixels in the local window, and cos() represents the calculation of the cosine similarity coefficient. They respectively represent the gradient vectors of the ith pixel and the jth pixel in the local window of the pixel X, and the gradient vector of each pixel is obtained by transforming the gradient of the depth map.

5. The method for three-dimensional reconstruction of urban buildings based on virtual modeling according to claim 1, characterized in that: Construct a plane with the optical center of view 1 and the optical center of view 2 and the point cloud x under view 1. Correspondingly, construct a plane with the optical center of view 1 and the optical center of view 2 and the point cloud under view 2. Construct a plane where point cloud x and point cloud is a pair of matched point clouds.

6. The method for three-dimensional reconstruction of urban buildings based on virtual modeling according to claim 1, characterized in that: The grayscale deviation corresponding calculation formula is: ; In the formula, Indicates the grayscale deviation of pixel Y in the grayscale image of view 1, Indicates the number of pixels in the local window, based on the actual number of pixels in the local window. Represents pixel pairs The absolute value of the grayscale difference between two pixels. Represents pixel pairs The absolute value of the grayscale difference between the pixels at the t-th position in the two local windows; Among them, the perspective 1 corresponds to the depth map a and the grayscale map a, and the perspective 2 corresponds to the depth map b and the grayscale map b. The two pixel points of the point cloud pairs corresponding to the perspective 1 and the perspective 2 in the depth map a and the depth map b are recorded as depth point pairs, and the two pixel points corresponding to the depth point pairs in the grayscale map a and the grayscale map b constitute pixel pairs.

7. The method for three-dimensional reconstruction of urban buildings based on virtual modeling according to claim 1, characterized in that: The corresponding calculation formula of the luminosity compensation weight is: ; In the formula, represents the photometric compensation weight of pixel k in the grayscale image of view 1, norm() represents the normalization function, represents the reflection standard difference of pixel k in the depth map of view 1, represents the depth homogeneity of pixel k in the depth map of view 1, Represents the grayscale deviation corresponding to pixel k in the grayscale image of view 1.

8. The method for three-dimensional reconstruction of urban buildings based on virtual modeling according to claim 1, characterized in that: The constructing of the loss function based on the photometric compensation weight further comprises: The difference between the depth map obtained by the deep learning model at each viewing angle and the depth value of the corresponding position in the label depth map is calculated. The differences obtained at all positions constitute the depth difference matrix at each viewing angle. The L1 norm of the product between the photometric compensation weight of each pixel point in the grayscale image of each viewing angle and the depth difference matrix of the corresponding viewing angle is used as the loss function.

9. A system for three-dimensional reconstruction of urban buildings based on virtual modeling, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Depth extraction method of light field image

    CN107135388A

  • Video image processing method based on ambient light standardization processing, medium and equipment

    CN118135162A