Method and device for automatically repairing photogrammetric model of unmanned aerial vehicle in urban environment

Through the improved YOLOv8 network and LaMa algorithm, combined with three-dimensional point projection counter-calculation technology, the ground occlusion problem in the drone photogrammetry model is automatically repaired, and the problems of grid topological structure deformation and texture mapping errors are solved, achieving efficient and real-time three-dimensional model repair.

CN120070267APending Publication Date: 2025-05-30Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510124202.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The drone photogrammetry model is caused by ground object occlusion in urban environments to deform grid topology and texture mapping errors. The existing manual editing methods are inefficient and cannot meet the real-time repair needs.

Method used

The improved YOLOv8 network structure is used for object detection, combined with LaMa algorithm and three-dimensional point projection inverse calculation technology, the target shadows are automatically extracted and the three-dimensional grid structure and texture mapping are repaired.

Benefits of technology

Automatic repair of drone photogrammetry model is realized, removing ground interference targets, restoring a high-quality realistic three-dimensional model, and improving repair efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070267A_ABST
    Figure CN120070267A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of photogrammetry and remote sensing three-dimensional reconstruction, and provides a method and a device for automatically repairing a photogrammetry model of an unmanned aerial vehicle in an urban environment. The method comprises the following steps of: firstly, accurately judging the position of a target by adopting an improved target detection method, and generating a corresponding binary mask at the same time; then adopting a local adaptive shadow extraction and discrimination strategy, ignoring other regions of the image, only performing shadow extraction on the periphery of the target, and performing target image restoration on the extracted target and the shadow thereof by adopting a LaMa algorithm; secondly, accurately deducing the position of a target three-dimensional grid point according to a target detection result of the unmanned aerial vehicle image through projection back calculation, automatically filtering other interference or discrete points and extracting boundary points of the target three-dimensional grid point to fit a plane where a target grid is located, so that the topological relation of the target grid is close to the same plane as a surrounding road grid; and finally, remapping the repaired texture surface patch and the mesh grid to realize automatic repair of the photogrammetry grid model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of photogrammetry and remote sensing three-dimensional reconstruction, and particularly relates to a method and device for automatically repairing an unmanned aerial vehicle photogrammetry model in an urban environment. Background Art

[0002] With the continuous development of oblique photogrammetry technology, the three-dimensional scene reconstruction technology based on aerial images can provide a three-dimensional model with geometric structure and rich texture for the ground, and has been widely applied in fields such as urban planning and traffic management. In recent years, due to the characteristics of high timeliness, high flexibility, and simple operation of the three-dimensional reconstruction of unmanned aerial vehicle oblique multi-view images, it has become a research hotspot in the field of photogrammetry and remote sensing. At the same time, the vigorous development of various three-dimensional reconstruction technologies also creates favorable conditions for the establishment of digital city three-dimensional models.

[0003] Currently, there are many relatively reliable methods for establishing a model from multi-view images. The main contents of the relatively mature multi-view three-dimensional reconstruction technology include steps such as camera calibration, dense matching, three-dimensional grid surface model generation, and texture mapping. However, when an unmanned aerial vehicle takes pictures in a real environment, due to the occlusion of ground objects, such as targets like vehicles and pedestrians, it will cause problems such as deformation of the grid topological structure and texture mapping error in the reconstructed scene. Especially in the road area, moving or stationary targets have a serious impact on the road grid quality. Although the unmanned aerial vehicle oblique measurement technology can obtain ground images from multiple perspectives, in a real environment, the occlusion of ground targets, especially in the road area, will inevitably damage the grid topological structure and texture, which is a problem that multi-view reconstruction cannot handle. And most of the processing methods for ground occlusion problems are to solve such problems through the editing of local region of interest grids or texture replacement. The manual selection and editing method has low efficiency and cannot meet the requirements of real-time repair of three-dimensional models. Summary of the Invention

[0004] Aiming at the problem that the existing unmanned aerial vehicle photogrammetry model is damaged to the grid topological structure and texture due to the occlusion of ground objects, and the editing or texture replacement of manually selecting and editing local region of interest networks cannot meet the requirements of real-time repair of three-dimensional models, the present invention proposes a method for automatically repairing an unmanned aerial vehicle photogrammetry model in an urban environment. First, targets are extracted from unmanned aerial vehicle images, and then the geometric structure and texture mapping of the three-dimensional grid where the targets are located are repaired in two aspects, so as to achieve the purpose of automatically removing ground interference targets and restoring a three-dimensional model with high-quality realism.

[0005] In a first aspect, the present invention provides a method for automatically repairing an unmanned aerial vehicle photogrammetry model in an urban environment, including:

[0006] Step 1: Input the pose - aware image sequence and the 3D mesh into the object detection network to detect the positions of the objects to be repaired in the image sequence, and generate a mask for the objects to be repaired based on the positions of the objects to be repaired; the pose - aware image sequence contains drone images with camera parameters.

[0007] Step 2: Further extract the shadow positions of the objects to be repaired based on the positions of the objects to be repaired, generate a shadow mask for the objects to be repaired, combine it with the mask for the objects to be repaired to form a mask to be repaired, and use the LaMa algorithm to repair the mask to be repaired to generate a target texture image.

[0008] Step 3: Use the projection back - calculation between 3D points and 2D image points according to the positions of the objects to be repaired to obtain the target mesh of the objects to be repaired in the 3D model, and flatten the target mesh to generate a mesh structure of the model to be repaired.

[0009] Step 4: Map the target texture image and the mesh structure of the model to be repaired to achieve automatic repair of the drone photogrammetry model.

[0010] Furthermore, the object detection network uses an improved YOLOv8 network structure.

[0011] Specifically, the Neck network of YOLOv8 is improved. The combination of the splicing layer and the up - sampling layer of the Neck network is replaced by an FFM module or an FEFM module, and a context - aware CAM module is added to the output of the Neck network; the FFM module is used to perform preliminary fusion processing on the features of different scales extracted by the Backbone network and the features of different scales extracted by the C2f layer, and the FEFM is used to further fuse features, effectively enhancing the feature fusion of different scales; the CAM module is used to integrate global information in the channel and spatial dimensions, and by multiplying with the input local features, highlight the importance of the target.

[0012] Correspondingly, the processing flow of the improved YOLOv8 network structure is as follows: Extract four features F 1 、F 2 、F 3 and F 4 from the Backbone network as inputs. First, F 1 、F 2 and F 3 undergo preliminary feature fusion through the FFM module, then undergo re - fusion after feature enhancement through the FEFM module, and finally, the CAM module captures valuable context information and outputs features F 1 ′、F 2 ′、F 3 ′and F 4 ′.

[0013] Further, the FFM module includes an upsampling layer, a convolutional layer, and a splicing layer;

[0014] Correspondingly, the FFM module includes an upsampling layer, a convolutional layer, and a splicing layer;

[0015] Correspondingly, the processing flow of the FFM module is as follows:

[0016] Determine that the inputs of the FFM module are the first feature and the second feature;

[0017] First, pass the first feature through the upsampling layer to obtain a third feature;

[0018] Input the third feature into the convolutional layer to obtain a fourth feature;

[0019] Splice the second feature, the third feature, and the fourth feature through the splicing layer to obtain the output of the FFM module.

[0020] Further, the FEFM module includes a convolutional layer, an FEM module, and a splicing layer. The FEM module includes four different branches. First, the four branches respectively pass through 1×1 convolutions. One of the branches is a residual structure, and the other three branches respectively pass through 3×3, 3×1, and 1×3 convolutions, and then perform dilated convolution operations with dilation rates of 3, 5, and 5 respectively. Finally, the four branches are spliced.

[0021] Further, the CAM module first performs global average pooling and max pooling on the input local features in the X direction and Y direction respectively, then performs a Concat operation on the generated features, reduces the dimension and activates through a 1×1 convolution, then performs a split operation along the spatial dimension, uses a 1×1 convolution to increase the dimension, combines with a sigmoid activation function, and finally multiplies with the input local features to obtain the features output by the CAM module.

[0022] Further, step 2 specifically includes:

[0023] Step 2 specifically includes:

[0024] Step 2.1: First, input the target position to be repaired, and expand w width outward with the coordinates of the target detection box of the target position to be repaired as the initial position to generate the region of interest of the target to be repaired and its shadow;

[0025] Step 2.2: Convert the image of the region of interest of the target to be repaired and its shadow from the RGB color space to the logarithmic domain, and then automatically extract the shadow region pixels through the local brightness threshold change of the region of interest of the target to be repaired and its shadow;

[0026] Step 2.3: Use the pixels in the shadow area as training features to construct a KNN classifier and classify the pixels in the shadow area; the KNN classifier uses the Euclidean distance as the distance metric and the nearest neighbor classification decision.

[0027] Step 2.4: Apply spatial filtering to the classified shadow image using a Gaussian kernel and binarize the shadow image to generate a target shadow mask to be repaired, which forms a mask to be repaired together with the target mask to be repaired.

[0028] Step 2.4: Use the LaMa model to repair the mask to be repaired to generate a target texture image.

[0029] Further, Step 2.3 further includes: passing the obtained mask to be repaired through an empirical threshold of the pixel area ratio of the target to the shadow to determine whether the shadow mask to be repaired in the mask to be repaired is a target shadow. If it is a target shadow, directly proceed to Step 2.4; if it is not a target shadow, then use the target mask to be repaired as the mask to be repaired and then proceed to Step 2.4.

[0030] Further, Step 3 specifically includes:

[0031] Input the target mask to be repaired, and after inverse projection calculation, obtain a preliminary target grid M, and based on the preliminary target network M, obtain a grid vertex set P = {P i};

[0032] Use the method of cloth simulation filtering on the grid vertex set P to remove other ground feature points outside the target to obtain a target grid vertex set P';

[0033] Based on the target grid vertex set P', set a search radius threshold and use a K-D tree for nearest neighbor search to obtain a target boundary point set P″ of the mask to be repaired in the 3D model;

[0034] Use the least squares method to fit the target boundary point set P″ to obtain the grid structure M' of the model to be repaired.

[0035] In a second aspect, the present invention provides an automatic repair device for an unmanned aerial vehicle photogrammetry model in an urban environment, including:

[0036] A target object detection module, configured to input an image sequence with poses and a 3D grid into a target detection network, detect the position of the target to be repaired in the image sequence, and generate a target mask to be repaired according to the position of the target to be repaired; the image sequence with poses includes unmanned aerial vehicle images with camera parameters.

[0037] The grid structure positioning module is used to further extract the shadow position of the target to be repaired based on the position of the target to be repaired, and generate a shadow mask of the target to be repaired, which forms a mask to be repaired with the mask of the target to be repaired. The LaMa algorithm is used to repair the mask to be repaired to generate a target texture image;

[0038] The shadow extraction module is used to perform projection back-calculation between three-dimensional points and two-dimensional image points according to the position of the target to be repaired, obtain the target grid of the target to be repaired in the three-dimensional model, and flatten the target grid to generate a grid structure of the model to be repaired;

[0039] The texture mapping module is used to map the target texture image and the grid structure of the model to be repaired to realize the automatic repair of the UAV photogrammetry model.

[0040] The beneficial effects of the present invention are as follows:

[0041] The present invention aims to solve the problem of grid occlusion in urban ground road photogrammetry, establish a correspondence between the interpretation on the two-dimensional image and the three-dimensional grid model, so as to accurately repair the target grid structure and texture. First, the topology of the three-dimensional target grid is automatically repaired by UAV target detection. Based on the adaptive extraction and discrimination of the target shadow in target detection, deep learning technology is used to complete the texture repair of UAV images to repair the texture mapping between the three-dimensional grid and the texture image. The entire process does not require manual editing and can be automatically processed. It is nested in the currently relatively mature MVS framework, which can effectively solve the problems of structural deformation and texture errors caused by ground target occlusion in the urban environment.

[0042] (1) The present invention provides an improved YOLOv8 network structure for target detection. The four different-scale features extracted by the backbone network are used as inputs, and after preliminary feature fusion by the FFM module, and then re-fusion after feature enhancement by the FEFM module. Finally, the valuable context information is captured by the context-aware CAM module to highlight the importance of the target and weaken the influence of the background on the target features, which can effectively improve the accuracy and efficiency of target detection.

[0043] (2) The present invention provides an algorithm for automatically extracting the target shadow of UAV images, which can remove the interference of shadows in the texture mapping process and provide a realistic texture for the model.

[0044] (3) The present invention uses projection back-calculation between three-dimensional points and two-dimensional image points to obtain three-dimensional grid points, and directly automatically flattens the vertices of the target grid. And only the target grid to be repaired is processed, avoiding waste of calculation and maintaining the authenticity of the surrounding grids. Description of the Drawings

[0045] Figure 1Schematic flowchart of an automatic repair method for an unmanned aerial vehicle photogrammetry model in an urban environment provided by an embodiment of the present invention;

[0046] Figure 2 Schematic structural diagram of an improved YOLOv8 network provided by an embodiment of the present invention;

[0047] Figure 3 Schematic structural diagram of an FFM module and an FEFM module provided by an embodiment of the present invention;

[0048] Figure 4 Schematic structural diagram of a CAM module provided by an embodiment of the present invention;

[0049] Figure 5 Schematic diagram of the detection results of the improved YOLOv8 network and different networks provided by an embodiment of the present invention;

[0050] Figure 6 Schematic structural diagram of a LaMa model provided by an embodiment of the present invention;

[0051] Figure 7 Schematic diagram of the comparison results of texture repair with and without shadow detection provided by an embodiment of the present invention;

[0052] Figure 8 Schematic flowchart of the process for generating a mesh structure of a model to be repaired provided by an embodiment of the present invention;

[0053] Figure 9 Schematic diagram of the results provided by an embodiment of the present invention;

[0054] Figure 10 Schematic diagram of the comparison of local model repair with and without shadows provided by an embodiment of the present invention;

[0055] Figure 11 Schematic diagram of the experimental result detection provided by an embodiment of the present invention. Detailed implementation manners

[0056] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0057] As Figure 1 shown, an automatic repair method for an unmanned aerial vehicle photogrammetry model in an urban environment provided by an embodiment of the present invention includes:

[0058] Step 1: Input the image sequence with pose and the 3D mesh into the target detection network to detect the position of the target to be repaired in the image sequence, and generate a mask for the target to be repaired according to the position of the target to be repaired; among them, the image sequence with pose contains UAV images with camera parameters;

[0059] Step 2: Further extract the shadow position of the target to be repaired based on the position of the target to be repaired, generate a mask for the shadow of the target to be repaired, and form a mask to be repaired with the mask of the target to be repaired. Use the LaMa algorithm to repair the mask to be repaired and generate a target texture image;

[0060] Step 3: Use the projection inverse calculation between 3D points and 2D image points according to the position of the target to be repaired to obtain the target mesh of the target to be repaired in the 3D model, and flatten the target mesh to generate a mesh structure of the model to be repaired;

[0061] Step 4: Map the target texture image and the mesh structure of the model to be repaired to realize the automatic repair of the UAV photogrammetry model.

[0062] Furthermore, as Figure 2 shown, the target detection network uses an improved YOLOv8 network structure; the Neck network of YOLOv8 is improved, and the combination of the splicing layer and the upsampling layer of the Neck network is replaced by an FFM module or an FEFM module, and a context-aware CAM module is added to the output of the Neck network; the FFM module is used to perform preliminary fusion processing on the features of different scales extracted by the Backbone network and the features of different scales extracted by the C2f layer, and the FEFM is used to further fuse the features, effectively enhancing the feature fusion of different scales; the CAM module is used to integrate global information in the channel and spatial dimensions, and highlight the importance of the target by multiplying with the input local features;

[0063] Correspondingly, the processing flow of the improved YOLOv8 network structure is: extract four features F 1 、F 2 、F 3 and F 4 of different scales in the Backbone network as inputs. First, F 1 、F 2 and F 3 undergo preliminary feature fusion by the FFM module, then undergo re-fusion after feature enhancement by the FEFM module, and finally the CAM module captures valuable context information and outputs features F 1 ′、F 2 ′、F 3 ′and F 4 ′.

[0064] Specifically, as Figure 3As shown, the FFM module includes an upsampling layer, a convolutional layer, and a splicing layer.

[0065] Correspondingly, the processing flow of the FFM module is as follows:

[0066] Determine that the inputs of the FFM module are the first feature and the second feature;

[0067] First, pass the first feature through the upsampling layer to obtain the third feature;

[0068] Input the third feature into the convolutional layer to obtain the fourth feature;

[0069] Splice the second feature, the third feature, and the fourth feature through the splicing layer to obtain the output of the FFM module.

[0070] Specifically, as Figure 3 shown, the FEFM module includes a convolutional layer, a FEM module, and a splicing layer, which can effectively enhance the feature fusion of different scales; among them, the FEM module aims to improve the receptive field of the network and adapt to multi-scale object detection, including four different branches. First, the four branches pass through 1×1 convolutions respectively. One of the branches is a residual structure, and the other three branches pass through 3×3, 3×1, and 1×3 convolutions respectively, so as to retain the features of the target at different scales through different convolution scales. Then, dilated convolutions with dilation rates of 3, 5, and 5 are performed respectively. Finally, the four branches are spliced together so that the extracted features contain more context information. Correspondingly, its calculation process can be expressed as:

[0071]

[0072] Y = Cat(X 1 , X 2 , X 3 , X 4 )

[0073] Among them, respectively represent conventional convolution operations with convolution kernel sizes of 1×1, 1×3, 3×1, and 3×3, and respectively represent dilated convolution operations with dilation rates of 3 and 5. Cat(Concat) represents the feature map splicing operation, and respectively represent the feature maps of the four branches obtained after the input feature map passes through the conventional convolution and the dilated convolution. Y represents the feature map after feature enhancement.

[0074] It can be understood that object detection is mainly achieved by using feature maps of different scales extracted by the Backbone network. These feature maps contain rich information such as the shape and position of the target. However, due to the limited features extracted by the backbone network, less semantic information, and limited receptive field, the feature enhancement FEFM module is constructed.

[0075] After the features are output by FFM and FEFM, the feature map has already focused on the local small target feature information at this time, and can represent the target features well. However, there is still a lack of relevance of context semantic information, resulting in the omission of many potential relationships between the target and the background, and a lack of the ability to distinguish between the target and the background. Therefore, the present invention constructs a context-aware CAM module, which is used after the feature fusion at four different scales. Referring to the practice of the channel and spatial attention CA module, global context information is established in the channel and spatial dimensions, thereby suppressing the influence of useless background on target detection.

[0076] Specifically, as Figure 4 shown, the CAM module first performs global average pooling and max pooling on the input local features in the X direction and Y direction respectively, then performs a Concat operation on the generated features, reduces the dimension and activates through a 1×1 convolution, then performs a split operation along the spatial dimension, and uses a 1×1 convolution to increase the dimension, and then combines with the sigmoid activation function. Finally, by multiplying with the input local features, the features output by the CAM module are obtained.

[0077] As Figure 5 shown, the embodiments of the present invention select the VisDrone-DET2019 dataset and the scene where the real captured data is located for comparing the detection results. Among them, Region 1 and Region 2 are the VisDrone-DET2019 dataset, and Region 3 and Region 4 are the real captured data.

[0078] Furthermore, the target is detected by the UAV target detection algorithm, and a mask is generated according to the bounding box of the detected target, and then image inpainting is performed. This is the situation in an ideal state. In the actual UAV shooting environment, the influence of sunlight causes shadows to appear on the target. The mask obtained only through target detection cannot include the position where the shadow is located, resulting in subsequent texture inpainting, and there will be a phenomenon that the vehicle is eliminated, but the shadow still exists, especially on the road. These shadow areas usually do not exceed the area occupied by the target, but due to the inability to determine the light direction, the specific direction of the target shadow cannot be judged. At the same time, in UAV images, in addition to the target shadow, there will also be shadows caused by various types of ground objects, which will all affect the discrimination of the target shadow. Therefore, on the basis of the above embodiments, a local adaptive shadow extraction strategy is proposed, which specifically includes:

[0079] Step 2.1: First, input the position of the target to be inpainted, and expand w widths outward with the coordinates of the target detection box at the position of the target to be inpainted as the initial position to generate the region of interest of the target to be inpainted and its shadow.

[0080] It can be understood that the target detection box coordinates at the target position to be repaired are extended by a width of w outward from the initial position to generate the region of interest of the target to be repaired and its shadow. This can avoid the influence of the shadows of other ground objects to the greatest extent.

[0081] Step 2.2: Convert the image of the region of interest of the target to be repaired and its shadow from the RGB color space to the logarithmic domain, which helps to enhance the contrast of the image and make the features of the shadow region more obvious. Then, automatically extract the shadow region pixels through the local brightness threshold change of the region of interest of the target to be repaired and its shadow.

[0082] Step 2.3: Construct a KNN classifier with the shadow region pixels as training features to classify the shadow region pixels; the KNN classifier uses the Euclidean distance as the distance metric and the nearest neighbor classification decision.

[0083] Step 2.4: Perform spatial filtering on the classified shadow image using a Gaussian kernel to smooth the image and reduce noise, and binarize the shadow image to generate a mask for the shadow of the target to be repaired, which together with the mask of the target to be repaired forms a mask to be repaired.

[0084] Step 2.4: Use the LaMa model to repair the mask to be repaired to generate a target texture image.

[0085] Specifically, after detecting and extracting the target and its shadow, it is further necessary to repair the pixels in the area where they are located to restore the complete texture information in the state without target occlusion. Although traditional methods can provide vivid textures for the repaired image, they often produce unrealistic images with repetitive patterns because they cannot capture high-level semantics. In addition, traditional methods cannot produce reasonable results for the repair tasks of complex regions with non-repetitive structures. Deep learning can greatly overcome the deficiencies of traditional methods by learning the semantic features of the input image and predicting the missing content based on these features. The resolution-robust large mask repair (LaMa) based on Fourier convolution has good repair effects on images with large missing areas, complex geometric structures, and high resolutions. The LaMa model introduces fast Fourier convolutions (FFCs) to obtain a larger receptive field. The network has the entire receptive field of the entire input image even in the shallow layer. The FFC not only improves the repair quality of the model but also reduces the number of model parameters. At the same time, the bias in the FFC makes the network have better generalization ability, and high-resolution repair results can be generated by training with low-resolution pictures. At the same time, a perceptual loss is used to further increase the receptive field. The network will first perform downsampling operations, then be processed by fast Fourier convolution, and finally upsampled to output the repaired image. The structure is as Figure 6As shown, due to its powerful performance in the field of image inpainting, it can directly process UAV images with a large width, and can meet the requirements of the present invention for repairing the target occlusion area.

[0086] Furthermore, step 2.3 further includes: judging whether the to-be-repaired shadow mask in the to-be-repaired mask is a target shadow through an empirical threshold of the pixel area ratio occupied by the target and the shadow. If it is a target shadow, directly proceed to step 2.4; if it is not a target shadow, then use the to-be-repaired target mask as the to-be-repaired mask and then proceed to step 2.4.

[0087] Since in the process of texture restoration, the pixels in the missing texture area rely on the surrounding pixels to be determined. Therefore, if the target shadow cannot be extracted and only the mask provided by target detection is relied on, the shadow will still exist in the model texture, which is a result we do not want to see. Figure 7 It can be seen that target detection is first performed on the original UAV image to obtain the initial position of the target. Then, based on the initial position of the target, the target shadow detection algorithm of the embodiment of the present invention is adopted, such as Figure 7 (e) and (f), the method of the present invention can better extract the shadow area where the target exists, providing a complete mask for texture restoration. The purpose of doing this is to remove the interference of the shadow in the texture mapping process and provide a realistic texture for the model.

[0088] Furthermore, as Figure 8 shown, step 3 specifically includes:

[0089] Input the to-be-repaired target mask, and after inverse projection calculation, obtain the preliminary target mesh M, and based on the preliminary target network M, obtain the mesh vertex set P = {P i};

[0090] Adopt the method of cloth simulation filtering for the mesh vertex set P to remove other ground object points outside the target, and obtain the target mesh vertex set P';

[0091] Specifically, the basic principle of the cloth simulation filtering algorithm is: first, flip the original point cloud, cover the simulated cloth above the flipped point cloud, and let it do free fall under the influence of gravity. If the cloth is soft enough, then the simulated cloth can completely cover the outer surface of the point cloud data, and its final shape is the digital surface model of the point cloud data. However, if the cloth has a certain hardness, then when doing the same free fall, the final shape of the cloth is the digital terrain model, which is the required ground points.

[0092] Set a search radius threshold based on the target mesh fixed-point set P', and use a K-D tree for nearest neighbor search to improve the search efficiency, and obtain the target boundary point set P″ of the to-be-repaired target mask in the three-dimensional model;

[0093] The least squares method is used to fit the target boundary point set P″ to obtain the mesh structure M′ of the model to be repaired.

[0094] The position of the target in the UAV image is determined according to the UAV target detection algorithm. According to the camera pose, the projection inverse calculation between the three-dimensional point and the two-dimensional image point is used to obtain the target grid vertex position of the target in the three-dimensional model from the two-dimensional coordinates detected by the target, and the corresponding three-dimensional grid vertex of the target is obtained. Taking a road vehicle as an example, the purpose is to restore it to the road grid without vehicles. Therefore, it is necessary to flatten the three-dimensional grid where the vehicle is located. The plane for target flattening is referenced by the surrounding road plane. If the surrounding road plane is additionally extracted, such as some modeling software like DP-Molder, by checking the road area and directly flattening the area by the fitting method, there is a problem: for a large-scale reconstructed model with a complex road environment, manual editing of the area of interest of the model road needs to be carried out multiple times through this software, which will inevitably result in high labor costs, low efficiency, and long cycle. Doing this is to directly flatten the target grid through the target detection result of the UAV image, and only process the target grid to be repaired. Therefore, the cloth simulation idea is adopted to filter out other discrete points outside the target. Since the boundary points of these points are often located on the outermost periphery, these boundary points are extracted, and the plane is fitted from the extracted three-dimensional points. This plane is regarded as the plane of the surrounding road grid of the target, avoiding waste of calculation and maintaining the authenticity of the surrounding grid.

[0095] Such as Figure 9 and Figure 10 shown, without manual editing, the repair of the overall texture and structure of the three-dimensional model is automatically completed. By Figure 10 comparing the local model repair with and without shadow extraction, Figure 11 (c)(f), traditional three-dimensional reconstruction will cause texture occlusion and artifacts. Only by repairing the texture based on the initial position provided by the target detection, the texture Figure 11 (d) is obtained. It can be seen that without the extraction of the target shadow, only the target vehicle is removed, but a large number of target vehicle shadows still remain on the texture. After the extraction of the target shadow of the present invention, as Figure 11 (e)(g), the target vehicle and its shadow are significantly removed, highlighting the importance of target shadow extraction.

[0096] The embodiment of the present invention also provides an automatic repair device for a UAV photogrammetry model in an urban environment, including:

[0097] A target detection module, configured to input the image sequence with pose and the three-dimensional grid into the target detection network, detect the position of the target to be repaired in the image sequence, and generate a mask of the target to be repaired according to the position of the target to be repaired; wherein, the image sequence with pose includes UAV images with camera parameters.

[0098] A grid structure positioning module is used to further extract the shadow position of the target to be repaired based on the position of the target to be repaired, generate a shadow mask of the target to be repaired, and form a mask to be repaired with the mask of the target to be repaired. The LaMa algorithm is used to repair the mask to be repaired to generate a target texture image;

[0099] A shadow extraction module is used to perform inverse projection calculation between three-dimensional points and two-dimensional image points according to the position of the target to be repaired, obtain the target grid of the target to be repaired in the three-dimensional model, and flatten the target grid to generate a grid structure of the model to be repaired;

[0100] A texture mapping module is used to map the target texture image and the grid structure of the model to be repaired to realize the automatic repair of the UAV photogrammetry model.

[0101] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for automatically repairing a UAV photogrammetry model in an urban environment, characterized in that: include: Step 1: Input the image sequence with pose and the three-dimensional grid into the target detection network, detect the position of the target to be repaired in the image sequence, and generate the mask of the target to be repaired according to the position of the target to be repaired; the image sequence with pose includes the drone image with camera parameters; Step 2: further extracting the shadow position of the target to be repaired based on the target position to be repaired, and generating a shadow mask of the target to be repaired, which is combined with the target mask to be repaired to form a mask to be repaired, and using the LaMa algorithm to repair the mask to be repaired to generate a target texture image; Step 3: According to the position of the target to be repaired, the projection back calculation between the three-dimensional point and the two-dimensional image point is used to obtain the target grid of the target to be repaired in the three-dimensional model, and the target grid is flattened to generate the grid structure of the model to be repaired; Step 4: Map the target texture image and the grid structure of the model to be repaired to achieve automatic repair of the UAV photogrammetry model.

2. The automatic repair method of drone photogrammetry model in urban environment according to claim 1 is characterized in that: The target detection network adopts an improved YOLOv8 network structure; Specifically, the Neck network of the YOLOv8 is improved, the combination of the splicing layer and the upsampling layer of the Neck network is replaced with an FFM module or a FEFM module, and a context-aware CAM module is added to the output of the Neck network; the FFM module is used to perform preliminary fusion processing on the features of different scales extracted by the Backbone network and the features of different scales extracted by the C2f layer, and the FEFM is used to further fuse the features to effectively enhance the fusion of features of different scales; the CAM module is used to integrate global information in the channel and spatial dimensions, and highlight the importance of the target by multiplying with the input local features; Correspondingly, the processing flow of the improved YOLOv8 network structure is as follows: extract four different scale features F1, F2, F3 and F4 in the Backbone network as input, firstly, F1, F2 and F3 are initially fused through the features of the FFM module, then they are fused again after feature enhancement by the FEFM module, and finally, the CAM module captures valuable context information and outputs features F′1, F′2, F′3 and F′4.

3. The automatic repair method of drone photogrammetry model in urban environment according to claim 2 is characterized in that: The FFM module includes an upsampling layer, a convolution layer and a splicing layer; Correspondingly, the FFM module processing flow is: Determine that the input of the FFM module is the first feature and the second feature; First, the first feature is passed through the upsampling layer to obtain a third feature; Inputting the third feature into the convolutional layer to obtain a fourth feature; The second feature, the third feature and the fourth feature are spliced ​​through the splicing layer to obtain the output of the FFM module.

4. The automatic repair method of drone photogrammetry model in urban environment according to claim 2 is characterized in that: The FEFM module includes a convolution layer, a FEM module and a splicing layer, wherein the FEM module includes four different branches. First, the four branches are respectively subjected to 1×1 convolution, one of which is a residual structure, and the other three branches are respectively subjected to 3×3, 3×1 and 1×3 convolutions, and then the dilated convolution operations with dilated rates of 3, 5 and 5 are respectively performed, and finally the four branches are spliced.

5. The automatic repair method of drone photogrammetry model in urban environment according to claim 2 is characterized in that: The CAM module first performs global average pooling and maximum pooling on the input local features in the X direction and Y direction respectively, then performs a Concat operation on the generated features, reduces the dimension and activates them through 1×1 convolution, then performs a split operation along the spatial dimension, and uses 1×1 convolution to increase the dimension, combined with the sigmoid activation function, and finally multiplies them with the input local features to obtain the features output by the CAM module.

6. The automatic repair method of drone photogrammetry model in urban environment according to claim 1 is characterized in that: Step 2 specifically includes: Step 2.1: First, input the position of the target to be repaired, and expand the target detection frame coordinates of the target to be repaired as the initial position to generate a region of interest of the target to be repaired and its shadow; Step 2.2: converting the image of the target to be repaired and its shadow region of interest from the RGB color space to the logarithmic domain, and then automatically extracting the shadow region pixels by changing the local brightness threshold of the target to be repaired and its shadow region of interest; Step 2.3: constructing a KNN classifier using the shadow area pixels as training features to classify the shadow area pixels; the KNN classifier uses Euclidean distance as a distance metric and adopts the nearest neighbor classification decision; Step 2.4: spatially filter the classified shadow image using a Gaussian kernel, and binarize the shadow image to generate a shadow mask of the target to be repaired, which is combined with the target mask to be repaired to form a mask to be repaired; Step 2.4: Use the LaMa model to repair the mask to be repaired and generate a target texture image.

7. The automatic repair method of drone photogrammetry model in urban environment according to claim 6 is characterized in that: The step 2.3 also includes: passing the obtained mask to be repaired through an empirical threshold of the pixel area ratio occupied by the target and the shadow to determine whether the shadow mask to be repaired in the mask to be repaired is a target shadow. If it is a target shadow, proceed directly to step 2.4; if it is not a target shadow, take the target mask to be repaired as the mask to be repaired, and then proceed to step 2.

4.

8. The automatic repair method of drone photogrammetry model in urban environment according to claim 1 is characterized in that: Step 3 specifically includes: Input the target mask to be repaired, obtain the preliminary target mesh M after projection back calculation, and obtain the mesh vertex set P based on the preliminary target mesh M = {P i }; The mesh vertex set P is subjected to cloth simulation filtering to remove other ground object points other than the target, so as to obtain the target mesh vertex set P'; A search radius threshold is set based on the target mesh vertex set P', and a KD tree is used to perform a neighbor search to obtain a target boundary point set P" of the target mask to be repaired in the three-dimensional model; The target boundary point set P″ is fitted by the least square method to obtain the grid structure M′ of the model to be repaired.

9. An automatic repair device for drone photogrammetry models in urban environments, characterized in that: include: The target detection module is used to input the image sequence with posture and the three-dimensional grid into the target detection network, detect the position of the target to be repaired in the image sequence, and generate a mask of the target to be repaired according to the position of the target to be repaired; the image sequence with posture includes the drone image with camera parameters; A grid structure positioning module is used to further extract the shadow position of the target to be repaired based on the target position to be repaired, and generate a shadow mask of the target to be repaired, which is combined with the mask of the target to be repaired to form a mask to be repaired, and the mask to be repaired is repaired by using a LaMa algorithm to generate a target texture image; A shadow extraction module is used to obtain a target mesh of the target to be repaired in the three-dimensional model by using projection back calculation between three-dimensional points and two-dimensional image points according to the position of the target to be repaired, and to flatten the target mesh to generate a mesh structure of the model to be repaired; The texture mapping module is used to map the target texture image and the grid structure of the model to be repaired to achieve automatic repair of the UAV photogrammetry model.

Citation Information

Cited By

  • Unmanned aerial vehicle image index construction method based on vector grating integration

    CN122388199A

  • An unmanned aerial vehicle image index construction method based on vector and raster integration

    CN122388199B