Vehicle Removal Method for Inclined 3D Models Based on Deep Learning

Through a deep learning-based method, using YOLOX and LAMA algorithms to automatically process vehicle removal and texture filling in the tilted three-dimensional model, the time-consuming and labor-consuming repair problems in the existing technology are solved and efficient automated processing is achieved.

CN116309680BActive Publication Date: 2025-07-25TIANJIN SURVEY DESIGN INST GRP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310154695.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-23
Publication Date
2025-07-25
Estimated Expiration
2043-02-23

AI Technical Summary

Technical Problem

In drone tilt photogrammetry, the vehicle geometric distortion and texture deformation in the three-dimensional model caused by vehicle position changes, the existing technology mainly relies on manual operations for repair in Adobe Photoshop, which is time-consuming and labor-consuming.

Method used

Using a deep learning-based method, the deep vehicle detection network YOLOX is used to locate the vehicle position, and then the vehicle erasing and texture filling are performed through the depth image repair algorithm LAMA to achieve automated processing.

Benefits of technology

Improves the efficiency of vehicle removal and texture filling, realizes end-to-end automated processing without manual intervention, and reduces operating time and labor intensity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116309680B_ABST
    Figure CN116309680B_ABST
Patent Text Reader

Abstract

The present invention provides a method for removing vehicles from an inclined three-dimensional model based on deep learning, which mainly includes three parts, namely data preprocessing, vehicle extraction, and vehicle removal. Beneficial effects of the present invention: It provides an algorithm for removing vehicles from an inclined three-dimensional model based on a deep convolutional neural network; aiming at the problem of vehicle deformation caused by vehicle movement in the inclined three-dimensional model, the present invention proposes a method for removing vehicles from an inclined three-dimensional model based on deep learning. First, the deep vehicle detection network YOLOX is used to detect vehicles in the three-dimensional model and locate the positions of the vehicles. Then, a deep image inpainting algorithm is used to erase the vehicles and automatically fill in the road texture. Compared with the manual vehicle removal algorithm based on Photoshop, the method proposed by the present invention can effectively improve the operation efficiency; the present invention is an end-to-end algorithm for removing vehicles from an inclined three-dimensional model without manual intervention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of optical remote sensing image processing, and in particular relates to a method for removing vehicles from an oblique three-dimensional model based on deep learning. Background Art

[0002] Oblique photography collects images from different angles such as vertical and oblique by carrying multiple sensors (usually five lenses) on the same flight platform, obtaining more complete and accurate information of ground objects, comprehensively perceiving complex scenes in a large range, high-precision, and high-clarity manner, and providing guarantees for the true effect of the real-scene three-dimensional scene and the mapping-level accuracy. Due to its convenient operation, flexibility, low cost, and being unrestricted by harsh environments and complex terrains, unmanned aerial vehicles (UAVs) can quickly complete large-area data collection and are widely used in oblique photogrammetry.

[0003] However, in UAV oblique photogrammetry, due to the change of vehicle positions in different image frames, geometric distortion and texture distortion of vehicles often occur in the obtained three-dimensional model, which seriously affects the accuracy and aesthetic effect of the oblique model. To address this problem, the currently adopted methods are mainly two: 1. Remove vehicles during the UAV image modeling process. For example, the Reconstruction Master software launched by Dashwise Intelligence Co., Ltd. determines whether a vehicle moves during the three-dimensional modeling process and eliminates the moving vehicles; 2. Remove deformed vehicles after three-dimensional modeling. This method is mainly used for repairing existing three-dimensional models. The commonly used method is to first flatten the deformed vehicle area, and then modify the texture image of this area in Adobe Photoshop software and manually erase the vehicle. For the repair of existing three-dimensional models, manual erasing and texture filling in Adobe Photoshop software are time-consuming and laborious. Therefore, how to automatically remove vehicles and fill textures using a deep convolutional neural network still needs further research. Summary of the Invention

[0004] In view of this, the present invention aims to propose a method for removing vehicles from an oblique three-dimensional model based on deep learning to address the problem of removing vehicles from an oblique three-dimensional model.

[0005] To achieve the above object, the technical solution of the present invention is realized as follows:

[0006] A method for removing vehicles from an oblique three-dimensional model based on deep learning includes the following steps:

[0007] S1. Preprocess training data;

[0008] S2. Use the preprocessed training data to train the deep vehicle detection network YOLOX;

[0009] S3. Use the deep vehicle detection network YOLOX to locate the vehicle, and then remove the vehicle based on the deep image restoration algorithm LAMA;

[0010] Further, the preprocessing of the training data in step S1 includes the following steps:

[0011] S11. Cut the large-scale 3D texture image into images of size 640×640 to obtain small images;

[0012] S12. Calculate the mean and variance of the small images and normalize the small images;

[0013] S13. Randomly select some small images as the training set, denoted as TrainData. Select some small images from TrainData for vehicle annotation, denoted as AnnoImg, and the remaining small images in TrainData are used for automatic annotation, denoted as UannoImg.

[0014] Further, the training of the deep vehicle detection network YOLOX using the preprocessed training data in step S2 includes the following steps:

[0015] S21. Use AnnoImg to train the deep vehicle detection network YOLOX of the UAV oblique 3D model through the training formula;

[0016] S22. Randomly select some small images from UannoImg, use the trained deep vehicle detection network YOLOX for vehicle detection, and write the vehicle detection results of each small image into a file to obtain a new preliminary vehicle annotation file, which is denoted as Anno;

[0017] S23. Visually edit Anno, correct the vehicle detection boxes and categories that are not detected correctly, and make it the annotation sample of UannoImg; input it into the deep vehicle detection network YOLOX to continue training and optimize the deep vehicle detection network;

[0018] S24. Repeat steps S22 - S23 until the vehicle detection results of the images in UannoImg meet the application requirements. The detection indicators of the application requirements are that the accuracy rate is greater than 90 and the recall rate is greater than 95;

[0019] S25. Output the optimized deep vehicle detection network YOLOX.

[0020] Further, the vehicle removal of the deep vehicle detection network in step S3 includes the following steps:

[0021] S31. Cut the large 3D texture image to be inferred into images of size 640×640 to obtain small 3D texture images;

[0022] S32. Use the depth vehicle detection network YOLOX obtained in step S25 to detect vehicles in the cut small 3D texture images, and obtain vehicle target detection frames, which are marked as CarBoundingBox;

[0023] S33. Generate a single-channel image of the same size as the small 3D texture image, which is marked as mask;

[0024] S34. Mark mask according to CarBoundingBox to obtain Dmask. In Dmask, the areas including vehicles are assigned 255, and the other area positions are assigned 0;

[0025] S35. Multiply Dmask and the small 3D texture image to obtain a color image mask with vehicle position markings, and input it into the depth image inpainting network LAMA for vehicle erasure and texture filling to obtain a small 3D texture image after removing vehicle targets;

[0026] S36. Merge the small 3D texture images after removing vehicle targets obtained in step S35 to synthesize a complete large 3D texture image without vehicles.

[0027] Further, the training formula in step S21 is:

[0028] L = Det(X, Y)

[0029]

[0030] where Det() represents the depth vehicle detection network, X represents the tilted 3D texture image, Y is the vehicle annotation result, L represents the network loss, θ is the depth vehicle detection network parameter, and γ is the learning rate.

[0031] Compared with the prior art, the method for removing vehicles from a tilted 3D model based on deep learning of the present invention has the following advantages:

[0032] The vehicle removal method for inclined 3D models based on deep learning described in the present invention provides an algorithm for removing vehicles from inclined 3D models based on a deep convolutional neural network. Aiming at the problem of vehicle deformation caused by vehicle movement in inclined 3D models, the present invention proposes a vehicle removal method for inclined 3D models based on deep learning. First, a deep object detection network is used to detect vehicles in the 3D model and locate the positions of the vehicles. Then, a deep image inpainting algorithm is used to erase the vehicles and automatically fill in the road textures. Compared with the manual vehicle removal algorithm based on Photoshop, the method proposed in the present invention can effectively improve the operation efficiency. The present invention is an end-to-end vehicle removal algorithm for inclined 3D models without manual intervention. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0034] Figure 1 is a schematic flowchart of steps S1 - S3 of the vehicle removal method for inclined 3D models based on deep learning according to an embodiment of the present invention;

[0035] Figure 2 is a schematic flowchart of step S2 of the vehicle removal method for inclined 3D models based on deep learning according to an embodiment of the present invention;

[0036] Figure 3 is a schematic diagram of cutting a large image into small images in the vehicle removal method for inclined 3D models based on deep learning according to an embodiment of the present invention;

[0037] Figure 4 is a schematic diagram of detecting vehicles in the 3D texture image in the vehicle removal method for inclined 3D models based on deep learning according to an embodiment of the present invention to obtain the vehicle target detection box CarBoundingBox;

[0038] Figure 5 is a schematic diagram of synthesizing the removed images into a complete large image in the vehicle removal method for inclined 3D models based on deep learning according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0040] In the description of the present invention, it should be understood that the orientation or positional relationships indicated by the terms "center", "longitudinal", "lateral", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. are based on the orientation or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more.

[0041] In the description of the present invention, it should be noted that unless otherwise clearly specified and defined, the terms "installed", "connected", "connected to" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood through specific situations.

[0042] The present invention will be described in detail below with reference to the drawings and in conjunction with embodiments.

[0043] As Figures 1 to 5 shown, for the method of removing vehicles from an inclined three-dimensional model based on deep learning, the present invention mainly includes three parts, namely data preprocessing, vehicle extraction, and vehicle removal. That is, for a given three-dimensional texture image, first, preprocessing such as normalization is performed on it, and then it is input into a vehicle detection convolutional neural network for vehicle extraction to determine the vehicle area and obtain a vehicle area map. Finally, the vehicle area map and the texture image are input into a depth image inpainting model to remove the vehicle and fill the texture.

[0044] In the present invention, the method of removing vehicles from an inclined three-dimensional model based on deep learning proposed by us has the following two remarkable features: First, the automatic detection and positioning of vehicles in the inclined three-dimensional model effectively reduces the time for manual inspection; second, the automatic vehicle removal and texture filling based on deep learning reduce the time for manual image editing.

[0045] To achieve the above object, the present invention includes the following steps:

[0046] S1: Preprocessing of training data. This step further includes:

[0047] 1.1 Cut the large 3D texture image into images of size 640×640 to obtain small images;

[0048] 1.2 Calculate the mean and variance of the small images and normalize the small images;

[0049] 1.3 Randomly select some small images as the training set, denoted as TrainData. Select some small images from TrainData for vehicle annotation, denoted as AnnoImg, and the remaining small images in TrainData are used for automated annotation, denoted as UannoImg.

[0050] S2: Train the deep vehicle detection network YOLOX using the preprocessed training data. This step further includes:

[0051] 2.1 Use the existing small amount of annotated data AnnoImg to train the deep vehicle detection network for the drone oblique 3D model. The formula is:

[0052] L = Det(X, Y)

[0053]

[0054] Where Det() represents the deep vehicle detection network, X represents the oblique 3D texture image, Y is the vehicle annotation result, L represents the network loss, θ is the parameter of the deep vehicle detection network, and γ is the learning rate.

[0055] 2.2 Randomly select some texture images from the large unannotated image set UannoImg, use the trained deep vehicle detection network YOLOX to detect vehicles, and write the vehicle detection results of each image into a file to obtain a new preliminary vehicle annotation file Anno;

[0056] 2.3 Visually edit Anno, correct the incorrectly detected vehicle detection boxes and categories, and make it the annotation sample of the unannotated image set UannoImg. Input it into the deep vehicle detection network YOLOX to continue training and optimize the deep vehicle detection network;

[0057] 2.4 Repeat steps 2.2 - 2.3 until the vehicle detection results of the images in the unannotated image set UannoImg meet the application requirements. The detection indicators of the application requirements are that the accuracy rate is greater than 90 and the recall rate is greater than 95;

[0058] 2.5 Output the optimized deep vehicle detection network YOLOX. The whole step is as Figure 2 shown;

[0059] S3: Use the deep vehicle detection network YOLOX to locate the vehicle, and then remove the vehicle based on the deep image inpainting algorithm LAMA. This step further includes:

[0060] S31. Cut the large-scale 3D texture image to be inferred into 640×640-sized images to obtain small-scale 3D texture images, as Figure 3 shown;

[0061] S32. Use the deep vehicle detection network YOLOX obtained in step S25 to detect vehicles in the cut small-scale 3D texture images to obtain the vehicle target detection box CarBoundingBox, as Figure 4 shown;

[0062] S33. Generate a single-channel image mask with the same size as the small-scale 3D texture image;

[0063] S34. Mark the mask according to the vehicle target detection box CarBoundingBox to obtain Dmask. In Dmask, the area containing the vehicle is assigned 255, and other areas are assigned 0;

[0064] S35. Multiply the Dmask image and the small-scale 3D texture image to obtain a color image mask with vehicle position markings, and input it into the deep image inpainting network LAMA for vehicle erasure and texture filling to obtain a small-scale 3D texture image after removing the vehicle target;

[0065] S36. Merge the small-scale 3D texture images obtained in step S35 after removing the vehicle target to synthesize a complete large-scale 3D texture image, as Figure 5 shown.

[0066] Advantages of the present invention:

[0067] (1) Provide a vehicle removal algorithm for inclined 3D models based on deep convolutional neural networks;

[0068] (2) Aiming at the vehicle deformation problem caused by vehicle movement in the inclined 3D model, the present invention proposes a method for removing vehicles from the inclined 3D model based on deep learning. First, use the deep vehicle detection network YOLOX to detect vehicles in the 3D model and locate the position of the vehicle, and then use the deep image inpainting algorithm (i.e., the deep image inpainting network LAMA) to erase the vehicle and automatically fill the road texture. Compared with the manual vehicle removal algorithm based on Photoshop, the method proposed by the present invention can effectively improve the operation efficiency;

[0069] (3) The present invention is an end-to-end vehicle removal algorithm for inclined 3D models without manual intervention.

[0070] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for removing vehicles from an inclined three-dimensional model based on deep learning, characterized in that: It includes the following steps: S1. Preprocess the training data; S2. Use the preprocessed training data to train the deep vehicle detection network YOLOX; S3. Use the deep vehicle detection network YOLOX to locate the vehicle, and then remove the vehicle based on the deep image inpainting algorithm LAMA; The preprocessing of the training data in step S1 includes the following steps: S11. Cut the large-scale three-dimensional texture image into images of size 640×640 to obtain small-scale images; S12. Calculate the mean and variance of the small-scale images and normalize the small-scale images; S13. Randomly select some small-scale images as the training set, denoted as TrainData. Select some small-scale images from TrainData for vehicle annotation, denoted as AnnoImg, and the remaining small-scale images in TrainData are used for automatic annotation, denoted as UannoImg; The use of the deep vehicle detection network YOLOX to locate the vehicle in step S3 and then remove the vehicle based on the deep image inpainting algorithm LAMA includes the following steps: S31. Cut the large-scale three-dimensional texture image to be inferred into images of size 640×640 to obtain small-scale three-dimensional texture images; S32. Use the deep vehicle detection network YOLOX to detect vehicles in the cut small-scale three-dimensional texture images to obtain vehicle target detection frames, and the vehicle target detection frames are marked as CarBoundingBox; S33. Generate a single-channel image with the same size as the small-scale three-dimensional texture image, and the single-channel image is marked as mask; S34. According to CarBoundingBox, mark mask to obtain Dmask. In Dmask, the area including the vehicle is assigned 255, and the other area positions are assigned 0; S35. Multiply Dmask and the small-scale three-dimensional texture image to obtain a color image mask with vehicle position marks, input it into the deep image inpainting network LAMA for vehicle erasure and texture filling to obtain a small-scale three-dimensional texture image after removing the vehicle target; S36. Merge the small-scale three-dimensional texture images obtained in step S35 after removing the vehicle target to synthesize a complete large-scale three-dimensional texture image without vehicles.

2. The method for removing vehicles from an inclined three-dimensional model based on deep learning according to claim 1, characterized in that: The use of the preprocessed training data to train the deep vehicle detection network YOLOX in step S2 includes the following steps: S21. Use AnnoImg to train the deep vehicle detection network YOLOX of the drone oblique three-dimensional model through the training formula; S22. Randomly select some small-scale images from UannoImg, use the trained deep vehicle detection network YOLOX to detect vehicles, write the vehicle detection results of each small-scale image into a file to obtain a new preliminary vehicle annotation file, and the new preliminary vehicle annotation file is denoted as Anno; S23. Visually edit Anno, correct the incorrectly detected vehicle detection frames and categories, and make it the annotation sample of UannoImg; input it into the deep vehicle detection network YOLOX for continuous training to optimize the deep vehicle detection network; S24. Repeat steps S22 - S23 until the vehicle detection results in UannoImg meet the application requirements. The detection metrics for the application requirements are that the accuracy rate is greater than 90 and the recall rate is greater than 95; S25. Output the optimized deep vehicle detection network YOLOX.

3. The method for removing inclined three-dimensional model vehicles based on deep learning according to claim 2, characterized in that: The training formula in step S21 is: ; ; Among them, Det() represents the deep vehicle detection network, X represents the tilted three-dimensional texture image, and Y is the vehicle annotation result. L represents the network loss. are the parameters of the deep vehicle detection network. is the learning rate.

4. An electronic device, comprising a processor and a memory communicatively connected to the processor and configured to store executable instructions of the processor, wherein: The processor is used to execute the method for removing vehicles from an inclined three - dimensional model based on deep learning according to any one of claims 1 - 3 above.

5. A server, characterized in that: It includes at least one processor and a memory communicatively connected to the processor. The memory stores instructions executable by the at least one processor. When the instructions are executed by the processor, the at least one processor executes the method for removing vehicles from an inclined three - dimensional model based on deep learning according to any one of claims 1 - 3.

6. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by the processor, it implements the method for removing vehicles from an inclined three - dimensional model based on deep learning according to any one of claims 1 - 3.

Citation Information

Patent Citations

  • Coal mine operation area calibration method based on computer vision

    CN111895931A

  • Traffic vehicle monocular positioning method, device and equipment and storage medium

    CN114255443A