An optimization method for air-ground collaborative real-world 3D modeling based on target detection and avoidance

By using drones and mobile phones to conduct collaborative surveys, combined with target detection algorithms, high-precision 3D modeling of densely built-up urban areas was achieved. This solved the problems of safety risks, incomplete sampling coverage, and high computational load that exist in traditional methods, improving the accuracy and detail of the model while reducing costs and computational load.

CN116030194BActive Publication Date: 2026-05-05CHINA CONSTR FIFTH ENG DIV CORP LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA CONSTR FIFTH ENG DIV CORP LTD
Filing Date
2023-02-06
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing technologies for 3D modeling of densely built-up urban areas suffer from problems such as flight safety risks, incomplete sampling coverage, low data accuracy, and high computational load. In particular, the 3D modeling results for cantilevered buildings and dense urban areas are not good, and dynamic interference targets in ground-based photographs can easily lead to a loss of model accuracy and detail.

Method used

A ground-air collaborative real-scene 3D modeling method based on target detection and avoidance is adopted. Through collaborative surveys using drones and mobile phones, target shooting scenes and areas are set. The YOLO target detection algorithm is used to identify and remove movable targets in ground shooting. Combined with EXIF ​​and POS information correction, unified aerial triangulation calculation of aerial and ground images is achieved.

Benefits of technology

It achieves high-precision, high-detail real-world 3D modeling under low-cost conditions, reducing computational load and modeling costs. It is particularly suitable for dense urban areas, improving the accuracy and detail of the model, and reducing redundant sampling and computational workload.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116030194B_ABST
    Figure CN116030194B_ABST
Patent Text Reader

Abstract

This invention discloses an optimization method for air-ground collaborative real-scene 3D modeling based on target detection and avoidance. In the preliminary research phase of real-scene 3D modeling, a collaborative research approach using drones and mobile phones is employed. This involves steps such as setting a unified target shooting scene, determining the target shooting area, acquiring ground-view images, acquiring sky-view images, and extracting and correcting EXIF ​​information to achieve unified aerial triangulation calculations and real-scene 3D modeling. Furthermore, a target detection algorithm is used to assist ground-based photographers in avoiding invalid ground targets, reducing the number of redundant ground-based sampling points, and simultaneously decreasing the workload of preliminary field research and subsequent aerial triangulation calculations. This invention's method enables high-precision, high-detail rapid real-scene 3D modeling under low-cost and high-efficiency acquisition conditions, coordinating aerial drone and ground mobile phone photo data into a unified aerial triangulation modeling process, achieving high-precision, high-detail rapid real-scene 3D modeling within limited cost and sampling conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of urban and rural building 3D modeling, and more specifically, it relates to an optimization method for air-ground collaborative real-scene 3D modeling based on target detection and avoidance. Background Technology

[0002] With the acceleration of urbanization, digital smart city technology is also developing rapidly. Ordinary flat maps can no longer meet the relevant needs of various industries. Compared with traditional two-dimensional maps, three-dimensional data maps have the advantages of being more realistic and having higher precision.

[0003] Currently, for buildings with special cantilevered structures, such as those with cantilevered roofs, eaves, and arcades, as well as densely built-up urban areas like urban villages, old residential areas, and old districts, traditional drone aerial photography faces limitations in flight safety, high costs, and even the risk of collisions. It can also lead to problems such as blurred details, rough accuracy, and even surface damage in 3D models, particularly regarding eaves and cantilevered areas. Because urban buildings are generally dense and there are many interfering targets, ground-based photography faces numerous limitations, including a limited number of ground path sampling points and interference from moving targets. From a sampling point perspective, the available routes for photography are limited, and it's impossible to achieve full coverage of the sampling area. Therefore, a certain degree of redundant sampling is needed at these limited points to ensure data quality, which increases the computational load.

[0004] While existing technologies exist that combine aerial and ground-based measurement methods for 3D building modeling, their effectiveness is often limited. This is because urban areas typically have dense buildings and numerous interfering targets. Ground-based photography faces limitations such as a limited number of sampling points along ground paths and interference from moving targets. From a sampling point perspective, the available routes for photography are limited, and it's impossible to achieve full coverage of the sampling area. Therefore, a certain degree of redundant sampling is needed at these limited points to ensure data quality, which also increases the computational load.

[0005] Existing aerial surveying techniques often require cameras to be fixed 360 degrees around buildings, ensuring there are no other moving or interfering targets around the buildings. However, this high-cost ground image acquisition is often not feasible for urban buildings with large populations. Ground-based photos of urban buildings are prone to containing other dynamic or interfering targets besides the buildings themselves. Furthermore, aerial and ground-acquired image data cannot be directly fused, resulting in a large computational load for data and model fusion. This leads to lower accuracy and detail in the 3D modeling of urban buildings, directly causing the loss of key details from the perspective of key people and key parts of buildings, which significantly affects designers' assessment of urban renewal issues.

[0006] For example, although Chinese patent CN202111288002.1 discloses a method for rapid local updates based on fine-grained real-scene modeling using mobile phone images, this method has certain limitations. Its technical process involves merging the aerial triangulation results from UAVs and those from mobile phones in post-processing, which is a method for updating 3D models through post-processing and modification. Although it improves the model, it still requires a lot of manual calculation and alignment modification, resulting in low efficiency due to manual repair. Furthermore, the same image feature points from different UAVs and mobile phone devices cannot participate in the same set of aerial triangulation calculations. The images collected by mobile phones also do not properly filter redundant targets when shooting urban buildings, all of which lead to increased measurement errors and aerial triangulation calculations.

[0007] Therefore, it is necessary to design an air-ground collaborative real-scene 3D modeling optimization method based on target detection and avoidance. Summary of the Invention

[0008] (I) Technical Issues

[0009] Based on the above-mentioned technical problems, this invention provides an optimization method for air-ground collaborative real-scene 3D modeling based on target detection and avoidance. This method can coordinate effective photo data captured by aerial drones and ground mobile devices such as mobile phones to a unified aerial triangulation modeling process under low-cost and high-efficiency acquisition conditions, thereby achieving high-precision, multi-detail real-scene 3D modeling under limited cost and limited sampling conditions. It also effectively reduces the computational load during image fusion and 3D modeling, and this method is particularly suitable for 3D modeling work in densely built-up urban areas.

[0010] This invention relates to a collaborative aerial-ground real-scene 3D modeling method that can quickly assist designers in completing traditional image surveys. It employs an object detection and avoidance mechanism to reduce the number of redundant ground samples, achieving high-precision, high-detail rapid real-scene 3D modeling under limited cost and sampling conditions. Furthermore, this method can be directly deployed in consumer-grade entry-level photographic acquisition and computational modeling products, significantly reducing the cost of real-scene 3D modeling.

[0011] (II) Technical Solution

[0012] This invention provides an optimization method for air-ground collaborative real-scene 3D modeling based on target detection and avoidance, the method comprising the following steps:

[0013] Step S1: Set the target shooting scene: Determine the shooting and acquisition window, which is the project requirement that both aerial and ground shooting must meet;

[0014] Step S2: Determine the target area to be photographed, and divide the photographing and acquisition task into two parts based on the target area: Step S3, ground shooting perspective work, and Step S4, aerial shooting perspective work.

[0015] Step S3: Ground shooting perspective operation: Set a set of shooting rules with the goal of efficiently acquiring building environment information from the ground shooting perspective. The purpose is to fill in dense obstacle areas and / or blind spots that are difficult for drones to shoot under the premise of safe operation. The image processing software in the portable mobile shooting device will detect as many movable targets as possible that cause ground interference, determine the number of movable targets in the captured image, and select photos that meet the basic shooting requirements. Then, proceed to step S5.

[0016] Step S4: Aerial shooting perspective work: Set a set of shooting rules with the goal of efficiently acquiring environmental information from the aerial shooting perspective. The purpose is to supplement the high-altitude perspective that is difficult to shoot from the ground under the premise of safe operation, so as to supplement the aerial and ground data required for air-ground collaborative aerial triangulation calculation. Then, proceed to step S6.

[0017] Step S5: Output and save all the ground-based image data with location information from step S3;

[0018] Step S6: Output and save all the aerial photography data with location information from step S4;

[0019] Step S7: Integrate the shooting data from the ground view in Step S5 and the aerial view in Step S6. By extracting EXIF ​​information, POS information and correcting POS information, integrate the photos collected from the aerial and ground shooting perspectives into an image group that can simultaneously satisfy the same aerial triangulation operation, providing a unified data foundation for air-ground collaborative aerial triangulation operation.

[0020] Step S8: Input the aerial and ground image groups that have been corrected and unified in step S7 into the aerial triangulation calculation.

[0021] Step S9: Generate a realistic 3D model of the building.

[0022] In another aspect, the present invention also discloses an air-ground collaborative real-scene 3D modeling optimization system based on target detection and avoidance, comprising:

[0023] At least one processor; and

[0024] At least one memory communicatively connected to the processor, wherein:

[0025] The memory stores program instructions that can be executed by the processor. The processor can call the program instructions to execute the air-ground collaborative real-scene 3D modeling optimization method based on target detection and avoidance as described above.

[0026] In another aspect, the present invention also discloses a non-transitory computer-readable storage medium storing computer instructions that cause the computer to execute the air-ground collaborative real-scene 3D modeling optimization method based on target detection and avoidance as described in any of the preceding claims.

[0027] (III) Beneficial Effects

[0028] (1) First, the three-dimensional modeling optimization method of the present invention adopts the method of simultaneous collaborative research of drones and mobile phones in the early research of real scene three-dimensional modeling. Through steps such as setting target shooting scene, determining target shooting area, collecting ground shooting perspective image, collecting sky perspective image, and extracting and correcting EXIF ​​information, it realizes unified aerial triangulation calculation of aerial and ground images and fast and high-precision modeling of real scene three-dimensional.

[0029] (2) Secondly, this invention also limits the shooting and acquisition windows such as time intervals, climate intervals, and spatial intervals, so that photos taken from the aerial shooting perspective and the ground shooting perspective can be taken in the same environment, ensuring that the background brightness of the aerial and ground images is consistent, so as to facilitate subsequent image fusion. In addition, in order to enable UAV photos and mobile phone photos to participate in the same set of aerial triangulation modeling calculations at the same time during image fusion, in addition to uniformly extracting EXIF ​​image information, the POS position is corrected and extracted based on the pose information of UAV and mobile phone, so as to ensure that the three-dimensional POS position information in the air and on the ground is located in the same coordinate system (latitude and longitude coordinates and / or arbitrary geographic coordinates) and the same altitude reference, providing a unified data foundation for air-ground collaborative aerial triangulation calculation, so as to improve the efficiency of subsequent three-dimensional modeling.

[0030] (3) In addition, the three-dimensional modeling optimization method of the present invention is particularly suitable for three-dimensional modeling of dense urban buildings. It can improve the sparsity effectiveness of photos taken by mobile devices such as ground mobile phones. By eliminating redundant photos of objects that interfere with the model, the ground photos can be better integrated with the aerial photos in three-dimensional data fusion. This greatly reduces unnecessary data processing and / or modeling calculation workload while ensuring the accuracy of the three-dimensional model. Attached Figure Description

[0031] Figure 1 This is a flowchart illustrating the steps of the air-ground collaborative real-scene 3D modeling optimization method based on target detection and avoidance in this invention.

[0032] Figure 2 This is a sample image of ground-based data acquisition prompts for target detection in an embodiment of the present invention.

[0033] Figure 3The diagrams shown are schematic diagrams of ground-based data acquisition scenarios in this embodiment of the invention, where (a) to (e) are schematic diagrams of moving targets under six different viewpoints.

[0034] Figure 4 This is a comparison chart of the aerial triangulation and modeling effects using sparse sampling and dense sampling based on target detection avoidance in embodiments of the present invention.

[0035] Figure 5 This is a 3D model of a small building created using conventional aerial flight method 1 in an embodiment of the present invention.

[0036] Figure 6 This is a 3D model of a small building modeled using the dense sampling-air-ground collaboration method 2 in an embodiment of the present invention.

[0037] Figure 7 This is a 3D model of a small building modeled using the sparse sampling-air-ground collaboration method 3 in an embodiment of the present invention. Detailed Implementation

[0038] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and are not intended to limit the scope of the invention.

[0039] With the advancement of urban renewal, the traditional data transmission method of delivering two-dimensional vector maps to planners and architects is no longer suitable. In vector map formats that only express two-dimensional wireframes, the actual 3D model calculated by aerial triangulation in traditional surveying methods is converted into a two-dimensional vector map through methods such as tracing and plotting by surveyors. This process loses most of the actual scene imagery information, retaining only spatial relationship information, resulting in a significant loss of data. Unlike new town construction projects, with the growth of urban renewal projects, designers need to frequently compare the differences between the actual scene and the renewal design. The demand for actual scene imagery information in renewal design is huge and growing rapidly, but it is currently rarely adequately met. There are three reasons for this: First, high-precision, large-area aerial triangulation is expensive and requires a professional team to operate, limiting the overall survey mobilization and collaboration. Second, the collection of key human perspectives and key parts of buildings, which are often most needed by designers, is often limited by flight obstacles and / or blind spots. In such densely built-up urban renewal areas, conducting this type of low-altitude aerial data collection poses considerable safety hazards. In three aspects, even after participating in certain shooting training and paying certain learning and training costs, ground surveying and data collection photographers still find it difficult to accurately reflect the updated design requirements during shooting and data collection. There are still some differences in understanding caused by different professions, backgrounds, and perspectives. Currently, excessive and redundant data collection is still used to cover up these differences.

[0040] Based on this, see Figure 1 As shown, this invention specifically proposes an air-ground collaborative real-scene 3D modeling optimization method based on target detection and avoidance, which includes the following steps:

[0041] Step S1: Set the target shooting scene: Determine the shooting acquisition window. The shooting acquisition window is the project requirement that both aerial and ground shooting must meet. The shooting acquisition window is determined with the goal of simultaneously meeting the aerial and ground data source requirements for aerial triangulation sampling.

[0042] Specifically, step S1 includes: the shooting and acquisition window includes a shooting time interval, a climate interval, and a spatial interval, wherein the definitions of each interval are as follows:

[0043] Time frame: Determine the working time conditions for air-ground collaboration and ensure that shooting conditions such as sunlight angle are controlled;

[0044] Climate range: Determine the common working climate conditions for air-ground collaboration to ensure that shooting conditions such as light intensity and visibility are controlled;

[0045] Spatial range: Determine the scope and conditions of the shared working space for air-ground collaboration, and ensure that shooting conditions such as latitude, longitude, and altitude are controlled.

[0046] Step S2: Determine the target area to be photographed. Based on the target area, divide the photographing and acquisition task into two parts: Step S3, ground shooting perspective work, and Step S4, aerial shooting perspective work.

[0047] Specifically, step S2 includes: the target area is expanded from the existing spatial range to a workable area where drones and mobile phones can safely capture and collect data, based on project requirements. Furthermore, the ground-based shooting perspective work in step S3 and the aerial shooting perspective work in step S4 can be performed simultaneously, or they can be performed separately, provided that the time, weather, and spatial timeframes are met.

[0048] Step S3: Ground Shooting Perspective Operation: Set a set of shooting rules with the goal of efficiently acquiring building environment information from the ground shooting perspective. The purpose is to fill in dense obstacle areas and / or blind spots that are difficult for drones to shoot under safe operating conditions. The image processing software in the portable mobile shooting device will detect as many movable targets as possible that cause ground interference, determine the number of movable targets in the captured images, and select photos that meet the basic shooting requirements. Then, proceed to step S5.

[0049] In step S3 above, considering that ground photography in existing technologies is limited by numerous moving ground targets (such as pedestrians, vehicles, bicycles, umbrellas, etc.), the effective information about the environmental scene in the field of view is very limited. This means that although the photographed scene can increase the detail of the model during real-world 3D construction, it also requires a large amount of redundant ground view sampling to ensure the model quality of air-to-ground collaborative computing. Excessive redundant ground view sampling further amplifies problems such as large data volume, computational complexity, and excessively large models. To address this, the shooting rules set based on portable mobile shooting devices such as mobile phones are designed to achieve high-efficiency acquisition of ground photos, supplement the ground data required for air-to-ground collaborative aerial triangulation calculations, and limit the data scale to avoid excessive interference from target data crowding out aerial triangulation computing resources.

[0050] Specifically, step S3 also includes:

[0051] S31: Determine the basic ground shooting angle: Determine the main subject to be shot from the ground perspective and the interfering objects that do not need to be shot, and determine the shooting angle, shooting height and shooting focal length.

[0052] Step S31 is to determine the shooting conditions. The main shooting objects include buildings, signs, shop signs, eaves, etc. Interference objects that do not need to be shot include moving targets such as pedestrians, cars, umbrellas, etc. The specific conditions can be determined according to the project and the site conditions. In addition, the portable mobile shooting device is preferably a mobile phone.

[0053] S32: Portable mobile shooting device identifies movable targets in the scene through target detection methods: The portable mobile shooting device is based on the YOLO (You Only Look Once) target detection algorithm, calls the trained model and weight file, constructs the target detection function (ObjectDetection()), and through the calculation of the detection backbone of YOLO's inferential sampling features, the hierarchical neck for widening the receptive field, and the predicted head for multi-scale output, the movable targets in the original scene are classified and detected according to type.

[0054] It is worth mentioning that the YOLO object detection algorithm, which is preferentially selected in this invention, has the advantages of small model size and fast computing speed, and can be directly deployed on portable mobile shooting devices commonly used on the ground, including but not limited to mobile phones and cameras. Currently, the most portable deployment method is to deploy the YOLO object detection algorithm in a mobile APP, because the graphics computing module (CPU and / or GPU unit), the location information acquisition module (GPS and / or Beidou), and the image capture module (camera) that are called for recognition can all be directly called on the mobile terminal at the same time, without the need for an external location acquisition device (POS) and / or RTK module, and also without the need for an external distributed computing device (graphics computing unit) for object detection. In addition, the object detection algorithm is based on object detection algorithms including but not limited to YOLO, SSD, R-CNN, etc., and can also be other lightweight object detection algorithms.

[0055] S33: Identify movable targets using the YOLO target detection algorithm, and label each type of movable target (label class_id, which corresponds to the category of the movable target), and calculate the position range parameters (x, y, w, h) and center point parameters (x+w / 2, y+h / 2) of the movable targets.

[0056] (x, y, w, h) represents the position and size parameters of the recognition box. First, (x, y) is the coordinate offset of the top-left corner of the recognition box relative to the top-left corner (0, 0) of the captured image (x is the rightward offset, y is the downward offset, in pixels); second, w and h are the width and height of the recognition box (starting from the top-left corner (x, y), w is the rightward offset from the starting point, and h is the downward offset from the starting point, in pixels); third, the point (x+w, y+h) defines the bottom-right corner of the recognition box (in pixels). Finally, based on the defined top-left corner (x, y) and bottom-right corner (x+w, y+h) of the recognition box, the recognition box for identifying the moving target can be drawn.

[0057] In addition, step S33 also includes step S331: for static environmental scenes where no moving targets are detected, no processing is performed, and the captured images are directly output as captured image sampling data, which serves as a data source for effective surface environment information in aerial triangulation.

[0058] Static environment scenes include buildings, trees, plaques, etc., which are the main subjects of the photography.

[0059] S34: For the movable target type (class_id) marked in step S33, draw the target center point based on the movable target center point parameters (x+w / 2, y+h / 2), and draw the target recognition box by superimposing the center point. The recognition box is jointly determined by the coordinate parameters (x, y) of the upper left corner point and the coordinate parameters (x+w, y+h) of the lower right corner point. The recognition box serves as a basis to assist the photographer in avoiding movable targets in the scene and to judge the shooting quality of the captured photos.

[0060] S35: Using superimposed recognition boxes as a discrimination tool, the photographer of the portable mobile shooting device can quickly identify and understand the environmental scene, assisting the photographer in determining whether the number of interfering movable targets in the scene meets the shooting condition requirements of step S31 (because some scenes cannot completely avoid the presence of movable targets in the image, the upper limit of the number of movable targets can be freely set to ensure that the number of movable targets is below the set threshold). If yes, proceed to the next step S36; if no, it means that there are too many interfering objects in the captured image, and the shooting condition requirements of step S31 cannot be met. In this case, it is necessary to return to step S31 to change the shooting scene and / or perspective (corresponding to...). Figure 1 (Step S351).

[0061] S36: When ground shooting meets the shooting requirements, images are captured using a portable mobile shooting device.

[0062] S37: The photographer determines whether the ground-based shooting is complete. If yes, proceed to step S5; otherwise, if the ground-based shooting is not complete, return to step S31 to change to the next scene and / or perspective after completing the current shooting scene and / or perspective (corresponding to...). Figure 1 (Step S371).

[0063] It should be noted that steps S31-S37 above are sparse sampling methods for ground images. This reduces the technical barrier to photography by avoiding target detection, better displaying architectural details, and lowering the learning and training costs for image acquisition. Furthermore, when implementing the shooting rules of steps S31-S37 on a mobile app, and since the app already defines the shooting rules and requirements, to reduce unnecessary travel for photographers, it is preferable to use crowdsourcing software to complete the ground shooting perspective. That is, the app can be used as an open crowdsourcing platform, allowing ordinary users to download and register the app to perform shooting tasks that meet the above rules and receive payment, thus enabling faster and better acquisition of ground images of dense building complexes.

[0064] Step S4: Aerial Photography Perspective Operation: Establish a set of shooting rules with the goal of efficiently acquiring environmental information from the aerial photography perspective. The aim is to supplement high-altitude perspectives that are difficult to capture from the ground under safe operating conditions, thereby supplementing the aerial and ground data required for air-to-ground collaborative aerial triangulation calculations. Then, proceed to step S6. Furthermore, the drone can be replaced with other portable aerial photography equipment.

[0065] Specifically, step S4 also includes:

[0066] S41: Determine the basic shooting angle from the air and ground: Determine the main subject of the ground perspective to be shot (such as buildings, sites, vegetation, etc.); at the same time, determine the camera's pan and head angle (Pitch, Yaw, Roll degree), shooting height, and shooting focal length.

[0067] S42: Plan the flight path of the drone in the air. Based on the project's shooting requirements, plan the flight path of the drone in the air and the shooting angle of the gimbal according to the methods such as "five-way flight", "grid flight", "zigzag flight" and / or "circling flight".

[0068] S43: When the target waypoint is reached and the aerial photography requirements are met, conduct drone photography.

[0069] S44: Determine whether aerial photography and data collection are complete based on the preset flight path. If yes, proceed to step S6; otherwise, if aerial photography is not complete, proceed to step S42 after completing photography at this waypoint and continue flying to the next waypoint (corresponding to...). Figure 1 Step S441 in the process.

[0070] Step S5: Output and save all the ground-based image data with location information from step S3. It should be noted that, during ground-based image acquisition, whether using a method of first performing a simplified detection and then capturing images, or using a method of first capturing densely and then filtering and simplifying photos, the principle is the same, and the protection received in this regard should be equal.

[0071] Step S6: Output and save all the aerial photography data with location information from step S4.

[0072] Step S7: Integrate the shooting data from the ground view in Step S5 and the aerial view in Step S6. By extracting EXIF ​​information, POS information and correcting POS information, integrate the photos collected from the aerial and ground shooting perspectives into an image group that can simultaneously satisfy the same aerial triangulation operation, providing a unified data foundation for air-ground collaborative aerial triangulation operation.

[0073] Specifically, step S7 also includes:

[0074] S71: Extracts EXIF ​​information from ground photos, including location information POS (latitude, longitude, and altitude), camera parameters such as lens model and lens focal length.

[0075] In addition, step S71 also includes step S711: extracting the EXIF ​​information of the aerial photo, including location information POS (latitude, longitude, altitude), gimbal shooting angle (Pitch, Yaw, Roll degree), shooting lens model, lens focal length and other camera parameters.

[0076] S72: Extract the POS location information from the EXIF ​​information of the ground photo, read the latitude and longitude (lon, lat) and height (Height) of the POS information of the ground photo, and correct it according to the project requirements to ensure that the POS location information on the ground and in the air is in the same coordinate system (latitude and longitude coordinates and / or any geographic coordinates) and the same height reference (relative to absolute altitude 0.000m and / or relative height relative to ground 0.000m).

[0077] In addition, step S72 also includes step S721: extracting the POS position information of the aerial photo from the EXIF ​​information, reading the latitude and longitude (lon, lat) and altitude (Height) of the POS information of the aerial ground photo, and correcting it according to project requirements to ensure that the POS position information of the aerial and ground are in the same coordinate system (latitude and longitude coordinates and / or arbitrary geographic coordinates) and the same altitude reference (relative to absolute altitude 0.000m and / or relative altitude 0.000m relative to the ground).

[0078] Step S8: Input the aerial and ground image groups that have been corrected and unified in step S7 into the aerial triangulation calculation.

[0079] Step S9: Generate a realistic 3D model of the building.

[0080] It is worth mentioning that the buildings modeled in the three-dimensional model of this invention are preferably dense buildings. Of course, this method can also be used to model other types of buildings in three dimensions.

[0081] In summary, to address the challenges of 3D modeling for densely packed buildings, this invention employs a collaborative survey method using both drones and mobile phones during the initial research phase of real-world 3D modeling. This involves steps such as setting a unified target shooting scene, determining the target shooting area, acquiring ground-view images, acquiring sky-view images, and extracting and correcting EXIF ​​information. This achieves unified aerial triangulation calculations and real-world 3D modeling. Furthermore, a target detection algorithm assists ground-based photographers in avoiding invalid ground-based targets, reducing the number of redundant ground-based sampling points, and simultaneously alleviating the workload of initial field research and subsequent aerial triangulation calculations.

[0082] Furthermore, based on steps S1-S9, the method of this invention can achieve deep fusion of aerial and ground photographs, unifying them into a single aerial triangulation measurement. This reduces the manual workload and difficulty of later model repair and maximizes the utilization of image feature point information from various angles to minimize measurement errors. For special architectural forms such as cantilevered structures, eaves, and arcades, this method has a good optimization effect, clearly defining spatial features such as under eaves and cantilevered structures, resulting in a clear structure and significantly improved local details of the model. For densely built-up urban renewal areas, mobile phone ground photography combined with drone photography is cost-effective, convenient for surveying and sampling, and improves the safety of both ground and aerial surveys (photographers do not need to climb to high places with their phones to avoid the risk of falling, and drones do not need to fly low to avoid the risk of collisions with buildings or trees). In addition, it reduces the technical difficulty and cost of ground photography and data collection, assists in efficient ground data collection, reduces the technical difficulty and training time of ground data collection work, and improves collection efficiency while reducing ineffective and inefficient redundant data collection work. While ensuring the quality and detail of the real-world 3D model data, the computational load of the overall aerial triangulation modeling is reduced, thus decreasing the data volume and computational cost of generating the real-world 3D model.

[0083] To illustrate the beneficial effects of the method of the present invention, the following also combines the following: Figures 2-7 The accompanying physical images and embodiments provide a detailed description of the 3D modeling optimization method for steps S1-S9 of the present invention and its advantages:

[0084] Example: Taking the air-ground collaborative 3D modeling of a small building as an example.

[0085] (1) Data Acquisition Optimization: During ground data acquisition in step S3, the YOLO target recognition algorithm deployed on the mobile phone provides real-time prompts. Figure 2 (Blue box and red dot) Determine whether movable, interfering targets exist in the current scene for ground-based data acquisition, and whether ground-based data acquisition is suitable. For cases involving movable, interfering targets, please refer to [link to relevant documentation]. Figure 3 , Figure 3 (a)-(f) Figure 6 Images of real-world scenes from various perspectives are automatically identified and marked as interfering ground photos on the phone.

[0086] (2) Efficiency Improvement: Regarding single ground sampling, comparing the results with and without sparse sampling (i.e., deleting ground photos containing movable targets and / or exceeding a set threshold for the number of movable target interference), the inventors found that reducing the number of ground samples from 16 to 10 reduced the workload of ground sampling and ground image calculation by 37.5%. Furthermore, in real-world modeling with single ground sampling, the single ground modeling effect after sparse sampling is almost unaffected. See details... Figure 4A comparison of the modeling effects.

[0087] (3) Test Comparison: The performance and results of the three methods were compared: traditional aerial flight (Method 1), dense sampling-air-ground collaboration using the present invention (Method 2, i.e., photos that do not meet the requirements of step S31 are not removed), and sparse sampling-air-ground collaboration using the present invention (Method 3, i.e., photos that do not meet the requirements of step S31 are removed). The basic test environment for comparing the performance and results of the three methods is as follows:

[0088] Hardware environment: CPU: Intel i9-10900k+ GPU: Nvidia 2080ti+ RAM: 32G; Software environment: Operating system: Windows 10-22H2+ Modeling software: Context Capture Master v4.4.5.33.

[0089] The comparison of computational efficiency and modeling results is shown in Table 1 below.

[0090] Table 1. Comparison of three sparse sampling methods using object detection avoidance and their modeling effects.

[0091]

[0092]

[0093] For details on the 3D modeling effects corresponding to methods 1-3 above, please refer to [link / reference]. Figure 5-7 Based on the above results, it can be concluded that: Method 1, as a traditional method, currently has a generally average modeling effect, but it performs well in areas where ground photography and data collection are not feasible. However, its aerial triangulation calculation and production modeling are relatively time-consuming. Method 2, which does not use step S3 of this invention, has the best modeling effect among all methods, but its corresponding ground data collection workload, computational modeling workload, and generated model file size also increase significantly. Its aerial triangulation calculation and production modeling are the most time-consuming compared to the other two methods. It performs well in non-dense building areas that require detailed modeling and where ground data collection can be freely carried out. Method 3, which uses steps S1-S9 of this invention, is a method that balances increased workload, model effect, and modeling efficiency. This method can use the YOLO algorithm for target detection and avoidance to complete sparse building image sampling with a limited increase in modeling workload, thereby effectively improving the modeling quality of fixed environmental parts such as buildings, signs, and vegetation, and achieving improved computational efficiency and reduced model file size to save system resources and improve the efficiency of post-processing of real-scene 3D modeling. Its aerial triangulation calculation and production modeling time are effectively controlled. It ensures the measurement details, accuracy, and recognizability of cantilevered spaces in special buildings such as under eaves, roof eaves, and arcades.

[0094] It is worth mentioning that the three-dimensional modeling method of the present invention can be converted into software program instructions, which can be implemented by a software analysis system including a processor and memory, or by computer instructions stored in a non-transitory computer-readable storage medium.

[0095] Finally, the method of this invention is merely a preferred embodiment and is not intended to limit the scope of protection of this invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for optimizing air-ground collaborative real-world 3D modeling based on target detection and avoidance, characterized in that, The method includes the following steps: Step S1: Set the target shooting scene: Determine the shooting and acquisition window, which is the project requirement that both aerial and ground shooting must meet; Step S2: Determine the target area to be photographed, and divide the photographing and acquisition task into two parts based on the target area: Step S3, ground shooting perspective work, and Step S4, aerial shooting perspective work. Step S3: Ground shooting perspective operation: Set a set of shooting rules with the goal of efficiently acquiring building environment information from the ground shooting perspective. The purpose is to fill in dense obstacle areas and / or aerial shooting blind spots that are difficult to shoot from the air under the premise of safe operation. The image processing software in the portable mobile shooting device will detect as many movable targets as possible that cause ground interference, determine the number of movable targets in the captured image, and select photos that meet the basic shooting requirements. Then, proceed to step S5. Step S4: Aerial shooting perspective work: Set a set of shooting rules with the goal of efficiently acquiring environmental information from the aerial shooting perspective. The purpose is to supplement the high-altitude perspective that is difficult to shoot from the ground under the premise of safe operation, so as to supplement the aerial and ground data required for air-ground collaborative aerial triangulation calculation. Then, proceed to step S6. Step S5: Output and save all the ground-based image data with location information from step S3; Step S6: Output and save all the aerial photography data with location information from step S4; Step S7: Integrate the shooting data from the ground view in Step S5 and the aerial view in Step S6. By extracting EXIF ​​information, POS information and correcting POS information, integrate the photos collected from the aerial and ground shooting perspectives into an image group that can simultaneously satisfy the same aerial triangulation operation, providing a unified data foundation for air-ground collaborative aerial triangulation operation. Step S8: Input the aerial and ground image groups that have been corrected and unified in step S7 into the aerial triangulation calculation. Step S9: Generate a realistic 3D model of the building.

2. The air-ground collaborative real-scene 3D modeling optimization method based on target detection and avoidance according to claim 1, characterized in that, Step S3 also includes: Step S31: Determine the basic ground shooting angle: Determine the main subject to be shot from the ground angle and the interfering objects that do not need to be shot, and determine the shooting angle, shooting height and shooting focal length. Step S32: The portable mobile shooting device identifies movable targets in the scene using a target detection method: Based on the YOLO target detection algorithm, the portable mobile shooting device calls the trained model and weight file to construct a target detection function. After calculating the detection backbone of YOLO's inference sampling features, the hierarchical neck for widening the receptive field, and the prediction head for multi-scale output, the movable targets in the original scene are classified and detected according to type. Step S33: Identify movable targets using the YOLO target detection algorithm, and label each type of movable target. Calculate the position range parameters (x, y, w, h) and center point parameters (x + w / 2, y + h / 2) of the movable targets. (x, y, w, h) represent the position and size parameters of the recognition box. (x, y) is the coordinate offset of the upper left corner of the recognition box relative to the upper left corner (0, 0) of the captured image, where x is the rightward offset and y is the downward offset. Next, w and h are the width and height of the recognition box, with the upper left corner (x, y) of the recognition box as the starting point. w is the rightward offset width from the starting point, and h is the downward offset height from the starting point. Step S34: For the movable target type marked in step S33, draw the target center point based on the movable target center point parameters (x + w / 2, y + h / 2), and draw the target recognition box by superimposing the center points. The recognition box is jointly determined by the coordinate parameters of the upper left corner point (x, y) and the coordinate parameters of the lower right corner point (x + w, y + h). The recognition box serves as a basis for assisting the photographer in avoiding movable targets in the scene and judging the shooting quality of the captured photos. Step S35: Using the superimposed recognition box as a discrimination tool, the photographer of the portable mobile shooting device can quickly identify and understand the environmental scene, and help the photographer determine whether the number of interfering movable targets in the scene can meet the shooting conditions required in step S31. If yes, proceed to the next step S36; if no, it means that the shooting conditions required in step S31 cannot be met, and it is necessary to return to step S31 to change the shooting scene and / or perspective. Step S36: When the ground shooting meets the shooting conditions, an image is captured using the portable mobile shooting device; Step S37: The photographer determines whether the shooting and acquisition of the ground has been completed. If so, the photographer jumps to step S5. If not, after completing the shooting scene and / or perspective, the photographer returns to step S31 to change to the next scene perspective.

3. The air-ground collaborative real-scene 3D modeling optimization method based on target detection and avoidance according to claim 2, characterized in that, The building is a dense building, the portable mobile shooting device is a mobile phone, and the APP application in the mobile phone is used to execute steps S31-S37, and the YOLO object detection algorithm is replaced with other lightweight object detection algorithms.

4. The air-ground collaborative real-scene 3D modeling optimization method based on target detection and avoidance according to claim 3, characterized in that, The mobile app in question is a crowdsourcing software.

5. The air-ground collaborative real-scene 3D modeling optimization method based on target detection and avoidance according to claim 4, characterized in that, Step S33 further includes step S331: for static environmental scenes where no movable targets are detected, no processing is performed, and the captured images are directly output as captured image sampling data, using the original image data as the effective data source of surface environment information in aerial triangulation.

6. The air-ground collaborative real-scene 3D modeling optimization method based on target detection and avoidance according to claim 5, characterized in that, The main subject to be photographed in step S31 includes buildings, signs, shop signs and / or eaves, and interfering objects that do not need to be photographed include moving pedestrians, cars and / or umbrellas.

7. The air-ground collaborative real-scene 3D modeling optimization method based on target detection and avoidance according to claim 1, characterized in that, Step S1 includes: the shooting and acquisition window includes the shooting time interval, climate interval, and spatial interval. Time frame: Determine the working time conditions for air-ground collaboration to ensure controlled shooting conditions at the angle of sunlight; Climate range: Determine the joint working climate conditions for air-ground collaboration to ensure that shooting conditions such as light intensity and visibility are controlled; Spatial range: Define the scope and conditions of the shared working space for air-ground collaboration, and ensure that shooting conditions with latitude, longitude and altitude restrictions are controlled.

8. The air-ground collaborative real-scene 3D modeling optimization method based on target detection and avoidance according to claim 1, characterized in that, Step S4 further includes: Step S41: Determine the basic aerial and ground shooting angle: Determine the main subject to be shot from the ground perspective; at the same time, determine the shooting lens and gimbal angle, shooting height, and shooting focal length. Step S42: Plan the flight path of the drone in the air. Based on the project's shooting requirements, plan the flight path of the drone and the shooting angle of the gimbal in the air according to the "five-way flight", "grid flight", "zigzag flight" and / or "circling flight" methods. Step S43: Once the target waypoint is reached and the aerial photography requirements are met, conduct drone photography. Step S44: Determine whether the aerial photography and data collection has been completed according to the preset flight path. If yes, proceed to step S6; otherwise, it means that the aerial photography has not been completed. After completing the photography at this waypoint, proceed to step S42 and continue to fly to the next waypoint for photography.

9. The air-ground collaborative real-scene 3D modeling optimization method based on target detection and avoidance according to claim 1, characterized in that, Step S7 further includes: Step S71: Extract the EXIF ​​information from the ground photo, including location information POS, lens model, lens focal length, and camera parameters; Step S72: Extract the POS position information from the EXIF ​​information of the ground photo, read the latitude, longitude and altitude of the POS information of the ground photo, and correct it according to the project requirements to ensure that the POS position information on the ground and in the air is in the same coordinate system and the same altitude reference.

10. The air-ground collaborative real-scene 3D modeling optimization method based on target detection and avoidance according to claim 9, characterized in that, Step S7 further includes: Step S711: Extract the EXIF ​​information of the aerial photo, including the position information POS, gimbal shooting angle, shooting lens model, lens focal length, and camera parameters; Step S721: Extract the POS position information from the EXIF ​​information of the aerial photo, read the latitude, longitude and altitude of the POS information of the aerial ground photo, and correct it according to the project requirements to ensure that the POS position information in the air and on the ground is in the same coordinate system and the same altitude reference.

Citation Information

Patent Citations

  • Local rapid updating method based on mobile phone image live-action refined modeling

    CN113963047A