Indoor and outdoor fine three-dimensional modeling method based on multi-modal data fusion
Through the indoor and outdoor fine 3D modeling method of multimodal data fusion and the use of various data collection methods and algorithms, the problems of low efficiency, poor accuracy and difficult integration of traditional modeling are solved, and efficient and accurate indoor and outdoor integrated 3D modeling is achieved to meet the application needs of complex scenarios.
Patent Information
- Application Number
- CN202510765046.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-10-17
AI Technical Summary
Traditional indoor and outdoor modeling methods are inefficient and costly, and it is difficult to achieve seamless docking and integrated expression of indoor and outdoor models. The accuracy of geometric registration and texture mapping is difficult to guarantee when fusing multimodal data, and semantic information extraction and classification are difficult.
By adopting a variety of data collection methods such as drone close-range photography, airborne laser scanning, ground supplementary photography, laser SLAM mobile scanning and panoramic cameras, combined with iterative closest point algorithm, generative adversarial network, SLAM pose graph correction and aerial triangulation adjustment solution, we can achieve spatiotemporal alignment of multi-source data, point cloud density completion and texture mapping, and construct a fine three-dimensional model integrating indoor and outdoor.
It improves the efficiency and accuracy of indoor and outdoor modeling, ensures data consistency and integrity, generates integrated indoor and outdoor 3D models with high geometric accuracy and rich texture details, and supports facility and equipment management and spatial analysis in complex scenarios.
Smart Images

Figure CN120807830A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of indoor and outdoor integrated modeling, and in particular to a multimodal data fusion method for indoor and outdoor fine three-dimensional modeling. Background Art
[0002] As the requirements for 3D model accuracy and details continue to increase in fields such as urban planning and architectural design, traditional indoor and outdoor modeling methods have gradually exposed many limitations.
[0003] In terms of indoor modeling, manual measurement is inefficient, labor-intensive, and difficult to obtain global information; although single-station 3D laser scanning can obtain high-precision point clouds, it requires multiple station settings, and data splicing is complex and costly; indoor photogrammetry is affected by occlusion and lighting, making it difficult to fully restore complex indoor environments.
[0004] When it comes to outdoor modeling, aerial photogrammetry lacks sufficient detail for depicting low-rise structures, while ground photogrammetry is inefficient and costly for modeling large areas. Furthermore, traditional indoor and outdoor modeling is often performed separately, making it difficult to seamlessly integrate and express geometric, texture, and semantic information between indoor and outdoor models. This leads to numerous inconveniences in practical applications, such as the inability to achieve integrated indoor and outdoor facility management and spatial analysis in smart urban management.
[0005] A multimodal data fusion method for indoor and outdoor detailed 3D modeling has emerged, leveraging a variety of data acquisition methods, including drone close-range photography, airborne laser scanning, supplementary ground photography, laser SLAM mobile scanning systems, and panoramic cameras. However, this method faces numerous challenges in practical application: UAV-collected data suffers from spatiotemporal inconsistencies and uneven point cloud density; indoor point cloud data contains gaps and missing data; geometric registration and texture mapping accuracy are difficult to ensure when fusing multi-source data; and issues such as model texture mapping, semantic information extraction and classification, and optimization and adjustment require urgent resolution. Summary of the Invention
[0006] In view of this, the purpose of the present invention is to develop a method for indoor and outdoor fine three-dimensional modeling that can effectively integrate multimodal data and overcome the difficulties of existing technologies. It can solve at least one technical problem mentioned in the background technology and has important practical significance and broad application prospects.
[0007] According to one aspect of the present invention, a method for indoor and outdoor fine 3D modeling using multimodal data fusion is provided, the method comprising:
[0008] Perform drone close-up photography, airborne laser scanning, and ground supplementary photography on the building in low-altitude areas to obtain the first photo dataset, the first laser point cloud dataset, and the POS dataset for the building's exterior.
[0009] The second photo data set and the second laser point cloud data set of the indoor building are obtained by continuously collecting at the indoor entrance through a laser SLAM mobile scanning system and a panoramic camera according to a preplanned route;
[0010] The multi-source data set is subjected to aerial triangulation adjustment calculation with the aid of the POS data set, and an indoor-outdoor integrated fine surface three-dimensional model is automatically constructed.
[0011] In the above technical solution, the method aims to obtain detailed information of indoor and outdoor buildings through various data acquisition means, and construct an indoor-outdoor integrated fine three-dimensional model by using advanced data processing and modeling technology. This process fully utilizes the advantages of different data acquisition methods, realizes the complementation and fusion of data, and provides rich data support for three-dimensional modeling.
[0012] Low-altitude area data acquisition: unmanned aerial vehicle close-range photography and airborne laser scanning are performed on the building in the low-altitude area to obtain a photo data set and a laser point cloud data set of the outdoor unmanned aerial vehicle accessible area of the building. Unmanned aerial vehicle photography can quickly obtain large-area image data, while laser scanning provides high-precision three-dimensional point cloud data, which together provide rich outdoor data for subsequent three-dimensional modeling.
[0013] Ground supplementary photography: hidden area textures that cannot be obtained by low-altitude unmanned aerial vehicle oblique photography are collected through ground supplementary photography at ground positions to obtain a ground photography data set carrying POS data, which is supplemented into the photo data set collected by the unmanned aerial vehicle, and the RTK positioning coordinates corresponding to each photo are matched through the shooting time of the ground photography. Ground supplementary photography can compensate for the blind area of unmanned aerial vehicle photography, ensure that the texture information of the hidden area is not lost, and the combination of POS data and RTK positioning coordinates improves the accuracy and traceability of the data.
[0014] Indoor data acquisition: indoor photo data sets and laser point cloud data sets are continuously collected at indoor entrances through a laser SLAM mobile scanning system and a panoramic camera according to a preplanned route. The laser SLAM system can realize real-time positioning and map construction in an indoor environment, and in combination with the panoramic camera, detailed texture and structure information of the indoor environment can be obtained to provide data support for indoor three-dimensional modeling.
[0015] Data fusion and modeling: the multi-source data sets collected above are subjected to joint aerial triangulation encryption adjustment in an aerial triangulation adjustment calculation software, and the aerial triangulation adjustment result is exported; based on the aerial triangulation adjustment result, an indoor-outdoor integrated surface three-dimensional model is constructed by using a modeling software. Aerial triangulation adjustment can geometrically correct and match the multi-source data, improve the consistency and accuracy of the data, and finally construct an indoor-outdoor integrated fine three-dimensional model.
[0016] The above method combines various data collection methods such as unmanned aerial photography, airborne laser scanning, ground supplementary photography, laser SLAM, and panoramic camera, fully utilizes their respective advantages, realizes data complementation and fusion, and improves the integrity and precision of the three-dimensional model. The method can simultaneously process indoor and outdoor data to construct an integrated indoor and outdoor three-dimensional model, meeting the demand for overall modeling of complex buildings.
[0017] In some embodiments, the building is subjected to unmanned aerial close-range photography, airborne laser scanning, ground supplementary photography, etc. in the low-altitude area to obtain a first photo data set, a first laser point cloud data set, and a POS data set of the building outdoors, which further includes:
[0018] If the first photo data set and the first laser point cloud data set are inconsistent in time and space, an iterative closest point algorithm is used to perform time and space registration, data synchronization is achieved through a timestamp sequence, spatial position deviation is corrected in combination with pre-calibration spatial coordinate parameters, and an outdoor registration data set consistent in time and space is obtained.
[0019] The third laser point cloud data set is extracted from the outdoor registration data set, the point cloud density is calculated, and if it is lower than a preset threshold, the missing data is inferred through a generative adversarial network combined with the geometric curvature and spatial distribution features of the adjacent point cloud to obtain a density-completed laser point cloud data set.
[0020] In the above technical solution, after unmanned aerial close-range photography and airborne laser scanning of the building in the low-altitude area, in order to solve the problem of time and space inconsistency that may occur in the collected photo data set and laser point cloud data set, as well as the problem of insufficient density of the laser point cloud data set, the methods of time and space registration and point cloud density completion are used. These steps ensure the consistency of data in time and space and improve the quality of the laser point cloud data set, providing a more accurate and complete data basis for subsequent three-dimensional modeling.
[0021] Time and space registration: First, check whether the photo data set and the laser point cloud data set are inconsistent in time and space. This may be caused by changes in the attitude of the unmanned aerial vehicle during flight, asynchronous shooting time, or mismatch between the scanning rate of the laser scanner and the movement speed of the unmanned aerial vehicle. If there is inconsistency in time and space, the iterative closest point (ICP) algorithm is used to perform time and space registration of the data. The ICP algorithm finds the closest point pairs between the two data sets, iteratively calculates and minimizes the distance between them, and thus realizes accurate registration of the data. At the same time, the timestamp sequence during data collection is used to synchronize the corresponding relationship of the photo and point cloud data in the time dimension. In combination with the pre-calibrated spatial coordinate parameters, the spatial position deviation of the data is corrected. This step uses known coordinate parameters to transform the data, so that data from different sources can be accurately aligned in space, and finally an outdoor registration data set consistent in time and space is obtained.
[0022] Point cloud density completion: Extract the laser point cloud dataset from the outdoor registration dataset and calculate its point cloud density. Point cloud density reflects the number of point cloud data points per unit area, which is one of the important indicators to measure the quality of point cloud data. If the calculated point cloud density is lower than the preset threshold, the point cloud density completion process is started. Low-density point cloud may result in a lack of detail or holes in the three-dimensional model, affecting the modeling quality. The generative adversarial network (GAN) is used to infer the missing data by combining the geometric curvature and spatial distribution characteristics of the neighboring point cloud. GAN can learn the distribution characteristics of the data and generate new data points. By analyzing the geometric shape and spatial distribution of the neighboring point cloud, the missing part is generated, which is consistent with the original data style and meets the geometric characteristics, so as to obtain the density completed laser point cloud dataset.
[0023] The above method solves the problem of time and space inconsistency that may occur in the collection process of multi-source data through the space-time registration step, ensures the accurate alignment and fusion of data, and provides a reliable data basis for subsequent modeling. The point cloud density completion step effectively improves the quality of the laser point cloud dataset, making the point cloud data more complete and uniform, which helps to generate more detailed and accurate three-dimensional models and reduces model defects caused by data loss.
[0024] In some embodiments, at the indoor entrance, the laser SLAM mobile scanning system and the panoramic camera are used to continuously collect along the pre-planned route to obtain a second photo dataset and a second laser point cloud dataset in the indoor building, which further includes:
[0025] The boundary features of the walls and obstacles are extracted from the initial indoor point cloud dataset, and the point cloud integrity is calculated. If it is lower than the preset threshold, the SLAM pose graph collected by the laser SLAM mobile scanning system is used to correct the point cloud coordinate deviation, and a complete indoor point cloud dataset is obtained.
[0026] In the above technical solution, the indoor data collection process not only obtains indoor photos and laser point cloud data through the laser SLAM mobile scanning system and the panoramic camera, but also pays special attention to the integrity of the point cloud data. By extracting the boundary features of the walls and obstacles and calculating the point cloud integrity, the point cloud coordinate deviation is corrected using the SLAM pose graph, and the indoor point cloud dataset is optimized, thereby improving the accuracy and quality of indoor three-dimensional modeling.
[0027] Indoor data collection: At the indoor entrance, the laser SLAM system is used for scanning. SLAM technology can locate and build a map in real time in an unknown environment, which is suitable for indoor environments where GPS signals are not available, and can provide high-precision point cloud data. In cooperation with the laser scanning system, indoor photo datasets are continuously collected along the pre-planned route to capture the texture and color information of the indoor environment.
[0028] Boundary feature extraction and point cloud integrity calculation: Extract the boundary features of walls and obstacles from the initial indoor point cloud dataset. These features are crucial for understanding the indoor space layout and modeling. Calculate the integrity of the extracted point cloud data by comparing the actual collected point cloud data with the expected amount of data that should be collected to assess the data's integrity and ensure its reliability.
[0029] Point cloud integrity optimization: If the calculated point cloud integrity is lower than the pre-set threshold, start the optimization process. This step uses the SLAM pose graph collected by the laser SLAM mobile scanning system to correct the point cloud coordinate deviation. The SLAM pose graph records the position and attitude information during the scanning process, which can be used to more accurately correct the spatial position of the point cloud data and improve the integrity of the point cloud data.
[0030] Optimized indoor point cloud dataset: Through the above steps, the integrity-optimized indoor point cloud dataset is obtained. The optimized dataset is more accurate in spatial position and has higher integrity, providing a more reliable data foundation for subsequent three-dimensional modeling.
[0031] The above method ensures the integrity and accuracy of indoor point cloud data through the integrity calculation and correction steps, providing high-quality data support for indoor three-dimensional modeling. The use of SLAM pose graph to correct point cloud coordinate deviation improves the spatial position accuracy of point cloud data, making the indoor model more accurate.
[0032] In some embodiments, multi-source dataset aerial triangulation adjustment is carried out with the assistance of POS dataset, and an integrated indoor and outdoor fine surface three-dimensional model is automatically constructed, including:
[0033] For the first and second laser point cloud datasets, color and texture features of visible light images in the first and second photo datasets are extracted, and after aerial triangulation adjustment processing, a least squares optimization algorithm is used to minimize the spatial error between point cloud coordinates and texture features for geometric registration, obtaining an integrated indoor and outdoor fine surface three-dimensional model with texture mapping.
[0034] In the above technical solution, the process aims to jointly perform aerial triangulation adjustment on the collected multi-source data (including the first and second photo datasets, ground supplementary photos, first and second laser point cloud data, POS data, etc.) to achieve accurate matching and calibration of the data. By fusing color and texture features in the laser point cloud dataset and the photo dataset, an integrated indoor and outdoor fine surface three-dimensional model with texture mapping is constructed.
[0035] Multi-source data input and aerial triangulation adjustment: Input the collected multi-source data sets (photo data sets, ground supplementary photos, laser point cloud data, POS data, etc.) into the aerial triangulation adjustment calculation software (such as ContextCapture software). Joint aerial triangulation adjustment is completed through matching calibration calculation. This process performs unified geometric correction and matching on data from different sources, ensuring consistency and fusion between data.
[0036] Fusion of laser point cloud data set and photo data set: For the laser point cloud data set, extract the color and texture features of the visible light image in the corresponding area of the photo data set. Use the least squares optimization algorithm to minimize the spatial error between point cloud coordinates and texture features for geometric registration. This step accurately aligns the point cloud data and photo texture through the optimization algorithm, ensuring the accuracy of texture mapping.
[0037] Fine surface three-dimensional model construction: Multi-source data set aerial triangulation adjustment is carried out with the assistance of POS data set, which can improve the accuracy and efficiency of the calculation. POS (Position and Orientation System) provides accurate camera position and attitude information, providing important initial conditions for aerial triangulation. Through aerial triangulation adjustment calculation, geometric correction and matching of multi-source data sets are performed to ensure data consistency and accuracy. Based on the optimized aerial triangulation results and fused data, professional modeling software (such as ContextCapture, Photoscan, etc.) is used to construct an indoor and outdoor integrated fine surface three-dimensional model. This model not only contains the geometric structure of the building, but also gives the model a realistic appearance through texture mapping, making it more realistic and practical.
[0038] The above method realizes high-precision fusion of multi-source data through joint aerial triangulation adjustment and least squares optimization algorithm, ensuring the geometric accuracy of the three-dimensional model and the accuracy of the texture details. The high-precision geometric information of the laser point cloud data and the rich color and texture information of the photo data are fused, making the generated three-dimensional model more outstanding in detail and better reflecting the real scene.
[0039] In some embodiments, multi-source data set aerial triangulation adjustment is carried out with the assistance of POS data set, and an indoor and outdoor integrated fine surface three-dimensional model is automatically constructed, which also includes:
[0040] During the aerial triangulation adjustment calculation process, the aerial triangulation matching results are evaluated based on the depth map.
[0041] In the above technical solution, during the adjustment and solution of aerial triangulation, the matching results of aerial triangulation are evaluated based on the depth map. The depth map provides depth information of the scene, which can help identify the distribution and consistency of matching points in space. Through depth map evaluation, errors or inconsistencies in aerial triangulation matching results can be detected, thereby improving the accuracy and reliability of matching.
[0042] In some embodiments, aerial triangulation adjustment and solution of multi-source data sets are carried out with the assistance of POS data sets, and an indoor and outdoor integrated fine surface three-dimensional model is automatically constructed. Subsequently, the following steps are included:
[0043] For the indoor and outdoor integrated surface mesh model, a weighted projection fusion algorithm of multi-view images is used to perform texture mapping with perspective coverage and lighting consistency as weights, generating an indoor and outdoor integrated surface mesh model with fine texture.
[0044] From the indoor and outdoor integrated surface mesh model with fine texture, semantic segmentation results are extracted, combined with boundary features, material properties and topological connection relationships, to generate a structured surface mesh model containing wall, floor and obstacle classification.
[0045] In the above technical solution, after the construction of the indoor and outdoor integrated fine surface three-dimensional model, in order to further improve the visual authenticity and semantic information richness of the model, a weighted projection fusion algorithm of multi-view images is used for texture mapping, and based on this, semantic segmentation results are extracted to generate a structured surface mesh model. This process makes the model not only have fine geometric structure and texture details, but also can provide clear semantic information for subsequent applications such as navigation and analysis.
[0046] Texture mapping optimization: Use image data obtained from multiple perspectives, combined with perspective coverage and lighting consistency as weights, to perform texture mapping. Perspective coverage ensures that all parts of the model are fully mapped, and lighting consistency ensures the coordination of textures under different lighting conditions. By weighted fusion of image information from multiple perspectives, an indoor and outdoor integrated surface mesh model with fine texture is generated. This step significantly improves the visual authenticity and detail performance of the model.
[0047] Semantic information extraction: Extract semantic segmentation results from the model with fine texture to identify different objects and regions in the model (such as walls, floors, obstacles, etc.). Combined with boundary features, material properties and topological connection relationships, the accuracy of semantic segmentation is enhanced. Boundary features help determine the boundaries of different objects, material properties distinguish the surfaces of different materials, and topological connection relationships maintain the structural integrity of the model.
[0048] Structured model generation: based on the semantic segmentation result, a structured surface mesh model containing wall, floor and obstacle classification is generated. This model not only contains geometric and texture information, but also adds semantic labels, which is convenient for subsequent application and analysis.
[0049] The above method uses fine texture mapping and semantic segmentation, and the model is more visually realistic and has rich semantic information, meeting the needs of more advanced applications. The integration of geometric, texture, semantic and other multi-dimensional information makes the model a comprehensive data carrier.
[0050] According to another aspect of the present application, an indoor and outdoor fine three-dimensional modeling system for multi-modal data fusion is provided, based on the above method, comprising:
[0051] An outdoor acquisition module is configured to perform unmanned aerial vehicle close-range photography, airborne laser scanning and ground supplementary photography on the building in the low-altitude area to obtain a first photo data set, a first laser point cloud data set and a POS data set of the building outdoors;
[0052] An indoor acquisition module is configured to continuously acquire a second photo data set and a second laser point cloud data set of the building indoors by using a laser SLAM mobile scanning system and a panoramic camera at the indoor entrance according to a pre-planned route;
[0053] A construction module is configured to perform aerial triangulation adjustment calculation on the multi-source data sets under the assistance of the POS data set, and automatically construct an integrated fine surface three-dimensional model of the indoor and outdoor.
[0054] According to still another aspect of the present application, an indoor and outdoor fine three-dimensional modeling device for multi-modal data fusion is provided, comprising:
[0055] At least one processor and a memory in communication with the at least one processor;
[0056] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the above method.
[0057] In the above technical solution, in order to better run and process the method, the above method is stored in the memory, and the processor is used to execute the stored method. It should be noted that the principle and effect of each step have been described above, and will not be expanded here.
[0058] According to still another aspect of the present application, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to implement the above method.
[0059] In the above technical solution, in order to better operate and use the method, the method is stored in a computer readable storage medium, and a processor is used to realize the method. It should be noted that the principle and effect of each step have been described above, and will not be expanded here. BRIEF DESCRIPTION OF DRAWINGS
[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.
[0061] Figure 1 is a flowchart of an embodiment of a multi-modal data fusion indoor and outdoor fine three-dimensional modeling method of the present application;
[0062] Figure 2 is a structural schematic diagram of an embodiment of a multi-modal data fusion indoor and outdoor fine three-dimensional modeling system of the present application. DETAILED DESCRIPTION
[0063] The present application will be further described in detail below in combination with the drawings and embodiments. It is particularly pointed out that the following embodiments are only used to illustrate the present application, but do not limit the scope of the present application. Similarly, the following embodiments are only some embodiments of the present application, not all embodiments, and all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0064] One of the embodiments
[0065] Please refer to Figure 1 A multi-modal data fusion indoor and outdoor fine three-dimensional modeling method, the method comprises:
[0066] S1, in the low altitude area, the unmanned aerial vehicle close-range photography, airborne laser scanning, ground supplementary photography and other ways are carried out on the building, and the first photo data set, the first laser point cloud data set and the POS data set of the building outdoor are obtained;
[0067] For example, a drone with a high-resolution camera and a high-precision laser scanner is selected. The high-resolution camera can clearly capture the details of the building's appearance, while the high-precision laser scanner can obtain the three-dimensional shape data of the building. At the same time, several backup batteries are prepared to ensure that the drone can complete long-term flight tasks. According to the layout of the building group and the surrounding environment, the flight area of the drone is divided into several blocks. The division of each area takes into account the distribution of buildings, flight safety, and the comprehensiveness of data collection. For example, the area with more concentrated buildings is divided into one block, and the open area is divided into another block. According to the height of the building, the complexity of the structure, and the required resolution, etc., the flight height, speed, and shooting interval of the drone are set. For example, for high and complex buildings, the flight height is set to a lower value, such as 30-50 meters, to obtain higher resolution images and point cloud data; at the same time, the flight speed is appropriately reduced, and the shooting interval is increased to ensure the integrity of data collection.
[0068] The drone flies according to the preset flight route and takes close-range photographs of the building in the low-altitude area. The camera automatically takes a large number of photos during flight, covering various angles and parts of the building. These photos include the front, side, back, top, and some detailed parts of the building, such as carvings, doors and windows, etc. For example, when taking a photograph of a high-rise building, the drone flies around the high-rise building from different heights and angles to obtain photographs of the overall appearance of the high-rise building and the details of the surface sign of the high-rise building. At the same time, the laser scanner continuously emits laser pulses during flight, receives reflected signals, and obtains three-dimensional point cloud data of the building. These point cloud data can accurately record the shape, size, and slight undulations of the building surface. The collected photos and point cloud data are downloaded to a computer for preliminary data preprocessing. For photo data, pictures that do not meet quality requirements such as blurring, overexposure, or underexposure are removed; point cloud data is filtered, denoised, and other processing is performed to remove some obvious noise points such as interference caused by birds and leaves.
[0069] For example, due to the limitations of low-altitude drone oblique photography, it is not possible to obtain texture information in some hidden areas (such as building gaps, shadow areas, and low structures). Therefore, ground supplementary photography is needed to collect texture data in these areas and integrate it into the photo data set collected by the drone.
[0070] Hidden area identification: Analyze the photo dataset and laser point cloud dataset collected by the UAV, identify hidden areas where insufficient texture information can be obtained. For example, narrow passages between two buildings in a commercial complex, shadow areas on the facade of a shopping mall, the bottom of a low-rise podium of a hotel, etc. Arrange technical personnel to carry out on-site investigation with ground photography equipment (such as high-definition cameras, tripods, etc.), confirm the specific location and range of hidden areas, and mark key points that need to be supplemented by photography.
[0071] Ground supplementary photography: Use ground photography equipment with high resolution and accurate positioning function, equipped with RTK (Real Time Kinematic) module to obtain accurate positioning coordinates. At the same time, ensure that the camera has good low-light performance to meet the shooting needs of shadow areas. According to the preset shooting plan, supplementary photography is carried out in hidden areas. For example, in the gap between buildings, use the wide-angle lens of the camera to shoot from different angles to ensure that the texture details of the buildings on both sides of the gap are collected; for the lower part of the low-rise structure, use the overhead shooting method to obtain the texture information of the bottom of the structure. Each photo records the shooting time and the corresponding positioning coordinates obtained through the RTK module. During the shooting process, the RTK module of the camera records the POS (Position and Orientation) data of each photo in real time, including longitude, latitude, height, pitch angle, roll angle and yaw angle, etc., to ensure that the geographical position and shooting attitude of the photo can be traced back.
[0072] Data integration and matching: Convert the photos collected by ground photography into the same format as the UAV photo dataset (such as JPEG or TIFF), and perform preliminary quality check on the photos to remove blurry, underexposed or overexposed photos. Use the shooting time stamp of ground photography to time-match the ground-collected photos with the UAV-collected photo dataset. For example, if the shooting time of a certain ground photo is close to the time when the UAV collects photos in the nearby area, it will be classified into the same time period dataset. Through the RTK positioning coordinates of each ground photo, it is matched with the positioning coordinates of the UAV-collected photos. For example, if the RTK coordinates of a certain ground photo are within a pre-set threshold range (such as 1 meter) from the coordinates of a certain photo collected by the UAV, they are considered to belong to the same area, and the ground photo is supplemented into the UAV photo dataset. The supplemented ground photos are fused with the UAV photo dataset to form a complete photo dataset covering all areas of the commercial complex, including hidden areas. At the same time, the POS data of the ground photos and the POS data of the UAV photos are integrated to ensure the positioning consistency of the entire dataset.
[0073] In this embodiment, S1, in the low-altitude area, the UAV close-range photography of the building, the airborne laser scanning, the ground supplementary photography, etc. obtain the first photo dataset, the first laser point cloud dataset, the POS dataset of the building outdoor, further comprising:
[0074] S12, if the first photo data set and the first laser point cloud data set exist spatial-temporal inconsistency, then use the iterative closest point algorithm to perform spatial-temporal registration, realize data synchronization through the timestamp sequence, correct the spatial position deviation combined with the pre-calibration spatial coordinate parameter, obtain the outdoor registration data set consistent in time and space;
[0075] S13, extract the third laser point cloud data set from the outdoor registration data set, calculate the point cloud density, if it is lower than the preset threshold, then infer the missing data through the generative adversarial network combined with the geometric curvature and spatial distribution characteristics of the adjacent point cloud, obtain the density completed laser point cloud data set.
[0076] For example, the photo data set and the laser point cloud data set of a commercial complex are collected by a drone, but it is found that there is a problem of spatial-temporal inconsistency between the photo data set and the laser point cloud data set in some areas of the commercial complex (such as the entrance square, parking lot, etc.). For example, part of the photos are taken at a low flight height and an angle of the drone, while the corresponding laser point cloud data is collected at a high flight height and a right angle of the drone, resulting in that the two cannot be directly matched in time and space.
[0077] Spatial-temporal registration operation: take the laser point cloud data set as the spatial reference, match the feature points in the photo data set with the nearest points in the laser point cloud data set. For example, in the entrance square area of the commercial complex, the algorithm finds the landmark building (such as arch, sculpture, etc.) at the entrance in the photo and the corresponding geometric feature point in the laser point cloud data, and gradually adjusts the spatial position of the photo data to make it consistent with the laser point cloud data in space. Use the timestamp sequence recorded when collecting the photo and the laser point cloud data to correspond the data in time sequence. For example, the photo taken at a certain time and the laser point cloud data scanned at the same time are matched through the timestamp to ensure their consistency in time. Combine the pre-set spatial coordinate parameters (such as the building coordinate system of the commercial complex, the take-off point coordinate of the drone, etc.) to further correct the spatial position of the data after spatial-temporal registration. For example, compare the registered data with the building plan of the commercial complex, find and correct the spatial position deviation, make the data accurately aligned in three-dimensional space, and finally obtain the outdoor registration data set consistent in time and space.
[0078] Point cloud density completion: Accurately extract the laser point cloud dataset from the outdoor registration dataset, which contains the three-dimensional point cloud information of each building structure in the commercial complex, such as the facade of the mall, the roof of the hotel, and the ground of the parking lot. Use professional software to calculate the point cloud density of the laser point cloud dataset. Divide the surrounding area of the commercial complex into a number of 1m x 1m x 1m cubic grids, count the number of point clouds in each grid, and calculate the overall point cloud density. For example, it is found that the point cloud density of the parking lot area around the commercial complex is 5 points per cubic meter, which is lower than the preset threshold (10 points per cubic meter), and cannot meet the needs of subsequent high-precision terrain analysis and model construction.
[0079] Generation of adversarial network construction and training: Collect existing high-density laser point cloud data around the commercial complex (such as high-precision scanning data at the entrance of the mall) as training samples, which contain rich geometric curvature and spatial distribution characteristics. Build a generative adversarial network, the generator uses U-Net architecture, and the discriminator uses PatchGAN architecture. The input of the generator is the geometric curvature and spatial distribution characteristics of the adjacent point cloud in the parking lot area, and the output is the completed point cloud data; the input of the discriminator is the real point cloud data and the generated point cloud data, and the output is the judgment of the data authenticity. Input the training samples into the generative adversarial network for training. During the training process, the generator generates missing data according to the input adjacent point cloud features, the discriminator distinguishes between generated data and real data, and feeds back to the generator, so that the generator continuously optimizes the generation effect. For example, through multiple iterations of training, the generator can gradually learn the distribution rules of the point cloud in the parking lot area, and generate completed data similar to the real point cloud. For parts of the parking lot area with point cloud density lower than the preset threshold, input the geometric curvature (such as the flatness of the ground, the shape of the parking space, etc.) and spatial distribution characteristics (such as the distribution trend of the point cloud in the horizontal and vertical directions, etc.) of the adjacent point cloud into the trained generative adversarial network, and infer the missing data through the network. For example, according to the point cloud features of the roads around the parking lot and the entrance of the mall, generate the missing point cloud data inside the parking lot, so as to obtain the density-completed laser point cloud dataset. At this time, the point cloud density of the parking lot area reaches 10 points per cubic meter.
[0080] S2, continuously collect the second photo dataset and the second laser point cloud dataset of the indoor building at the indoor entrance by the laser SLAM mobile scanning system and the panoramic camera according to the pre-planned route;
[0081] For example, in order to obtain detailed data of the indoor entrance and its surrounding area, a laser SLAM mobile scanning system and a panoramic camera are used for data collection.
[0082] A set of reliable and positioning-capable laser SLAM mobile scanning system is selected, which includes a high-precision laser scanner and an inertial measurement unit (IMU) for real-time acquisition of indoor environment laser point cloud data and attitude information during movement. At the same time, a high-resolution panoramic camera is equipped, which can shoot 360-degree panoramic photos, and the resolution of the camera is not less than 20 million pixels to ensure that the collected photos have enough details. In addition, a stable mobile platform (such as a robot chassis equipped with a laser SLAM system and a panoramic camera) is prepared for indoor movement and data collection. According to the layout map of the indoor entrance of the commercial complex, a reasonable data collection route is planned in advance. This route should cover all key areas at the entrance, including the entrance passage, the lobby, the elevator hall, the stairwell, the main corridor, and the main commercial area connected to the entrance, etc. For example, starting from the entrance gate, along the main passage, through the lobby, elevator hall, stairwell, and then into the main passage and store entrance area of the mall connected to the entrance, to ensure comprehensive collection of indoor data at the entrance and its surroundings.
[0083] Data collection operation: At the starting point of the indoor entrance (such as the open area outside the entrance gate), the laser SLAM mobile scanning system and the panoramic camera are initialized and calibrated. Turn on the equipment, set parameters such as laser scanning frequency, camera exposure time, etc., and perform zero-point calibration and attitude calibration to ensure the accuracy and stability of the equipment during the collection process. Start the robot chassis to move along the pre-planned route, and the laser SLAM system and panoramic camera start data collection simultaneously. During the movement, the laser scanner emits laser beams at a certain frequency (such as 10-20 times per second) to scan the surrounding environment and obtain three-dimensional point cloud data of the indoor space; the panoramic camera takes panoramic photos at a set time interval (such as every 2-5 seconds) to record the texture information of the indoor scene. For example, when moving in the entrance passage, the laser SLAM system continuously scans the walls, floor, and ceiling on both sides of the passage to obtain three-dimensional point clouds of these parts; at the same time, the panoramic camera takes panoramic photos inside the passage, including decorative, signage, store signage, and other texture information. The collected laser point cloud data and panoramic photos are transmitted in real time to the data storage device (such as a notebook computer or a mobile hard disk) on the mobile platform through the data line, and stored and managed according to certain file formats and naming rules. For example, store the laser point cloud data as.pcap format files, and the panoramic photos as.jpg or.png format files, while recording the collection time, location information, and device parameters of each set of data as metadata.
[0084] Data quality inspection and supplementary collection: During and after the collection process, timely quality inspection is conducted on the obtained data. Through the visualization software, the integrity and accuracy of the laser point cloud data are checked to see if there are any data loss, noise interference or insufficient coverage areas; at the same time, the panoramic photos are browsed to check if the clarity, brightness and color saturation of the photos meet the requirements. For the data quality problems found, such as insufficient coverage of the laser point cloud data at a corner or unclearness of the panoramic photo in a certain area, timely supplementary collection is arranged to ensure that the collected data is reliable and complete.
[0085] Data preprocessing: The collected laser point cloud data is preliminarily processed, including filtering and denoising, coordinate system conversion and data splicing operations. The noise points in the point cloud data are filtered out, such as abnormal points caused by mirror reflection or device error; the laser point cloud data is converted from the device coordinate system to the unified indoor world coordinate system; the multiple sections of point cloud data collected along the route are spliced to generate a continuous and complete indoor three-dimensional point cloud model. For example, the point cloud data of the entrance hall area is spliced using the software provided by the laser SLAM system or professional point cloud processing software (such as CloudCompare) to form a complete three-dimensional point cloud model of the hall. The panoramic photos are simply preprocessed, such as color correction, image splicing and format conversion, etc. The color balance, contrast and brightness of the photos are adjusted to make the colors more realistic and natural; the multiple panoramic photos taken (such as photos of adjacent areas taken at different positions) are spliced to generate a larger range of panoramic view; the photos are converted to formats suitable for subsequent three-dimensional modeling and texture mapping, such as.exr or.tiff formats. For example, the panoramic photos taken at different positions in the entrance passage are spliced into a complete passage panoramic view using panoramic splicing software for subsequent texture mapping.
[0086] In this embodiment, S2, a second photo data set and a second laser point cloud data set of the indoor building are obtained by continuously collecting at the indoor entrance according to a pre-planned route using a laser SLAM mobile scanning system and a panoramic camera, which further comprises:
[0087] S21, the boundary features of the walls and obstacles are extracted from the indoor initial point cloud data set, the point cloud integrity is calculated, and if it is lower than a preset threshold, the SLAM pose graph collected by the laser SLAM mobile scanning system is used to correct the point cloud coordinate deviation to obtain an indoor point cloud data set with optimized integrity.
[0088] For example, due to reasons such as shaking, occlusion and environmental complexity during the movement of the device, the boundary features of the walls and obstacles in some areas are missing or incomplete, and the point cloud data needs to be optimized to improve its integrity and usability.
[0089] Data preprocessing and preliminary inspection: Import the collected indoor initial point cloud dataset into professional point cloud processing software (such as CloudCompare, MATLAB, or Python's Open3D library). Use visualization tools to view the overall distribution and details of the point cloud data, and preliminarily understand whether there are obvious missing or incomplete areas in the boundary features of the walls and obstacles. For example, in the entrance hall area, it is found that the point cloud data of part of the wall edge and the elevator button area is relatively sparse, and there is obvious boundary missing. Perform simple cleaning on the initial point cloud data, including removing noise points (such as outliers, and error points caused by mirror reflection) and filtering. Statistical filtering or radius filtering methods can be used to remove abnormal points that are too far or too dense from surrounding points, improving data quality.
[0090] Wall and obstacle boundary feature extraction: Use point cloud segmentation algorithms based on geometric features (such as region growing, clustering analysis, or machine learning methods) to segment the indoor point cloud data into different regions, and extract walls, obstacles (such as columns, elevators, stair railings, etc.), and other ground and ceiling parts. For example, according to the planar characteristics of the walls and the shape features of the obstacles, the walls in the entrance hall area and the obstacles in the elevator area are segmented from the overall point cloud. Apply boundary detection algorithms (such as methods based on curvature, normal vector, or neighborhood analysis) to extract the boundary feature points of the walls and obstacles. Calculate the curvature and normal vector of each point to identify the feature points at the boundary, which usually have large curvature changes or different normal vector directions. For example, for the junction between the wall and the door frame, the points on the boundary line are accurately extracted through curvature and normal vector analysis.
[0091] Point cloud integrity calculation: Determine the evaluation indicators of point cloud integrity, such as the continuity, coverage rate, and missing rate of the walls and obstacle boundaries. For example, calculate the ratio of the distance between boundary feature points to the expected boundary length to evaluate the continuity of the boundary; calculate the ratio of the actual collected wall and obstacle point cloud area to the theoretical area to evaluate the coverage rate. According to the selected evaluation indicators, calculate the point cloud integrity of the walls and obstacles in the indoor initial point cloud dataset. For example, it is found that the continuity of the wall boundary in the entrance hall area is 85%, and the coverage rate is only 70%, which is lower than the preset integrity threshold (such as continuity 90%, coverage rate 80%), indicating that the point cloud data in this area needs to be optimized.
[0092] Correcting Point Cloud Coordinate Deviations Using a SLAM Pose Graph: A SLAM pose graph is acquired from the laser SLAM mobile scanning system during the acquisition process. This pose graph records the device's position and posture information (such as position coordinates and rotation angles) as it moves indoors, reflecting the device's motion trajectory and posture changes. The SLAM pose graph is compared and analyzed with the initial point cloud dataset to identify areas of point cloud coordinate error caused by device pose deviations. For example, when the device passes through the foyer area, uneven ground and human interference cause deviations in the position and posture recorded in the pose graph, resulting in offsets in the point cloud coordinates of walls and obstacles. Using the accurate pose information in the SLAM pose graph, the point cloud coordinates in the initial point cloud dataset are corrected. Using coordinate transformation algorithms (such as rotation matrix and translation vector calculations), the point cloud data is converted from the device coordinate system to the real indoor world coordinate system to correct for point cloud coordinate deviations. For example, based on the accurate position and posture of the foyer area recorded in the pose graph, the point clouds of walls and obstacles in the corresponding area are transformed to more accurately align with the actual geometric positions.
[0093] Completeness optimization: Based on the corrected point cloud coordinates and extracted boundary features, missing wall and obstacle boundary features are completed. Interpolation algorithms (such as linear interpolation and spline interpolation) or geometric model-based repair methods (such as plane fitting and cylindrical fitting) can be used to infer the point cloud data of the missing part based on the existing boundary feature points and the distribution of the surrounding point clouds. For example, for the missing part of the wall edge, the adjacent boundary feature points and the wall plane fitting results are used to generate the missing edge point cloud through linear interpolation. The completed point cloud data is fused with the initial point cloud dataset to form an indoor point cloud dataset with optimized completeness. The fused data is further smoothed and optimized to remove redundant and abrupt points, thereby improving the overall quality and visualization of the point cloud. For example, the moving least squares (MLS) method is used to smooth the point cloud to make the boundary between the wall and the obstacle smoother and more natural.
[0094] S3. With the assistance of POS datasets, aerial triangulation adjustment and solution of multi-source datasets are carried out, and a fine surface 3D model integrating indoor and outdoor elements is automatically constructed.
[0095] In this embodiment, S3, with the assistance of the POS dataset, performs aerial triangulation adjustment of the multi-source dataset and automatically constructs a fine surface three-dimensional model that integrates indoor and outdoor features, including:
[0096] S31, for the first and second laser point cloud data sets, extract the color and texture features of the visible light images in the first and second photo data sets, after the aerial triangulation adjustment processing, adopt the least square optimization algorithm to minimize the spatial error of the point cloud coordinates and the texture features as the target for geometric registration, and obtain the indoor and outdoor integrated fine surface three-dimensional model with texture mapping.
[0097] For example, the outdoor photo data set, the ground supplementary photo data set, the indoor photo data set and the laser point cloud data set have been collected by means of unmanned aerial photography, ground supplementary photography, laser SLAM mobile scanning system and panoramic camera respectively. Now it is necessary to comprehensively process these data, extract relevant features and perform geometric registration to construct an indoor and outdoor integrated fine surface three-dimensional model.
[0098] The outdoor photo data set collected by the unmanned aerial vehicle, the photo data set collected by the ground supplementary photography, and the indoor photo data set collected by the laser SLAM mobile scanning system and the panoramic camera are integrated together. The initial position and attitude information of the unmanned aerial vehicle aerial image provided by the POS (position and attitude system) is used to carry out aerial triangulation adjustment calculation on the multi-source data set (including unmanned aerial vehicle close-range photography, ground supplementary photography and airborne laser scanning, and the second photo data set and the second laser point cloud data set). In this process, the characteristics and error factors of various data are comprehensively considered, and the mathematical model and algorithm are established to optimize and adjust the interior orientation elements, exterior orientation elements of the image and the position and attitude of the point cloud data, so that all the data are geometrically configured in the best way in a unified coordinate system. At the same time, the system will automatically construct an indoor and outdoor integrated fine surface three-dimensional model according to the calculated results. For example, in the outdoor part of the building, the building facade, roof structure and surrounding terrain are accurately restored according to the image and laser point cloud data; in the indoor part, the layout, interior decoration and three-dimensional form of various facilities of the room are constructed in detail according to the indoor collected image data and laser point cloud data, realizing seamless splicing and transition of indoor and outdoor models, and forming a complete, fine and highly realistic city area three-dimensional model.
[0099] Data preparation and preprocessing: the outdoor photo data set collected by the unmanned aerial vehicle, the photo data set collected by the ground supplementary photography, and the indoor photo data set collected by the laser SLAM mobile scanning system and the panoramic camera are integrated together. Ensure that all photo data carries positioning data, and the laser point cloud data set has matching time with the corresponding photo data. The photo data of different sources is unified in format, such as converted to TIFF format, and the laser point cloud data is preprocessed to filter out noise points and outliers to ensure data quality.
[0100] Feature extraction: Use image processing software (such as MATLAB, Python's OpenCV library) to extract color and texture features from visible light images in the photo dataset. Color features can be obtained by calculating the color histogram, color moment, etc. of the image; texture features can be extracted using the gray level co-occurrence matrix (GLCM), LBP algorithm, etc. For example, for the photos of the facade of a commercial complex, extract the color information (such as red brick walls, white windows) and texture features (such as the rough texture of the brick wall, the smooth texture of the glass) of the surface material. For the laser point cloud dataset, use point cloud processing software (such as CloudCompare, Python's Open3D library) to extract geometric features such as point cloud coordinates, normal vectors, and curvature. These features will be used for subsequent geometric registration.
[0101] Aerial triangulation adjustment processing: Construct an aerial triangulation model to determine the position and attitude parameters of the camera by matching the same points between different photos. Use software (such as Photomodeler, ContextCapture) to automatically identify feature points in photos and perform feature matching to establish the initial connection between the three-dimensional point cloud and the photo. According to the initial parameters obtained by aerial triangulation, perform adjustment processing to optimize the position and attitude parameters of the camera, reduce measurement errors and observation errors. Use least squares optimization algorithm to make the adjusted parameters better conform to the actual geometric relationship and improve the model accuracy.
[0102] Geometric registration and texture mapping: Take the laser point cloud dataset as the reference, and register the color and texture features of the visible light images extracted from the photo dataset with the point cloud coordinates. Use the least squares optimization algorithm to minimize the spatial error between the point cloud coordinates and the texture features as the objective function, adjust the position and attitude of the texture features to make them accurately aligned with the point cloud coordinates. For example, for the wall at the entrance of a commercial complex, register the color and texture features extracted from the wall surface with the coordinates of the wall in the laser point cloud dataset, so that the texture can be accurately mapped to the surface of the wall. Map the registered texture features to the three-dimensional model of the laser point cloud dataset to generate a fine surface three-dimensional model with texture mapping. Use texture mapping algorithms (such as interpolation algorithm, projection algorithm) to assign color and texture information of the image to the surface of the point cloud model, so that the model not only has geometric shape, but also has realistic visual appearance. For example, in the interior of a commercial complex, map the texture of floor to the ground point cloud, map the texture of ceiling to the top point cloud, and map the wall texture to the wall point cloud, so that the three-dimensional model can realistically reflect the actual scene of the indoor.
[0103] In this embodiment, S3, under the assistance of POS data set, multi-source data set aerial triangulation adjustment is carried out, and an indoor and outdoor integrated fine surface three-dimensional model is automatically constructed, which further includes:
[0104] S32, in the process of aerial triangulation adjustment, the aerial triangulation matching result is evaluated based on the depth map.
[0105] For example, in the process of aerial triangulation adjustment, the system will evaluate the aerial triangulation matching result based on the depth map. The depth map is obtained by processing the image, which can reflect the depth information of each pixel point in the scene. The system compares and analyzes the homonymic points obtained by aerial triangulation with the depth information in the depth map. For example, for the corners or complex structure parts of some buildings, if the homonymic points matched by aerial triangulation show large depth difference on the depth map, it may mean that the matching has errors. The system will automatically adjust and optimize the position and parameters of the matching points according to the evaluation result of the depth map, so as to improve the accuracy and reliability of aerial triangulation matching, so as to ensure that the three-dimensional model constructed later more accurately and truly reflects the geometric characteristics of the actual scene.
[0106] In this embodiment, S3, under the assistance of POS data set, multi-source data set aerial triangulation adjustment is carried out, and an indoor and outdoor integrated fine surface three-dimensional model is automatically constructed, which further includes:
[0107] S4, for the indoor and outdoor integrated surface grid model, a weighted projection fusion algorithm of multi-view images is used to perform texture mapping with view coverage rate and lighting consistency as weights, to generate an indoor and outdoor integrated surface grid model with fine texture;
[0108] For example, an indoor and outdoor integrated surface grid model has been constructed. In order to further improve the visual effect and detail performance of the model, it is necessary to use the weighted projection fusion algorithm of multi-view images for texture mapping.
[0109] Multi-view image preparation: Collect high-quality images taken from multiple angles to ensure coverage of all key indoor and outdoor areas. These images should have high resolution and good lighting conditions to capture rich texture details. For example, use a drone to collect images of the building facade, and use ground photography equipment to collect images of indoor entrances, shopping mall interiors, office areas, etc. Preprocess the collected images, including color correction, image enhancement, etc. to improve the quality and consistency of the images. Color correction can eliminate color differences between different images, and image enhancement can improve the contrast and clarity of the images.
[0110] View coverage and lighting consistency evaluation: For each view image, calculate its coverage on the three-dimensional model surface. The view coverage can be evaluated by projecting the image onto the three-dimensional model and calculating the proportion of the projected area to the total surface area of the model. For example, for an image of a building facade, its projected area covers 20% of the facade. Analyze the lighting conditions of each view image to evaluate its lighting consistency. This can be evaluated by calculating features such as lighting intensity, direction, and shadow distribution. Images with good lighting consistency usually have uniform lighting distribution and fewer shadow obstructions.
[0111] Weight calculation: Assign weights to each view image based on view coverage and lighting consistency. A weighted average method can be used to normalize view coverage and lighting consistency to the [0, 1] interval, and then assign different weight coefficients (e.g., view coverage weight coefficient is 0.6, lighting consistency weight coefficient is 0.4) to calculate the comprehensive weight. For example, for an image with view coverage 0.2 and lighting consistency 0.8, its comprehensive weight is 0.6 x 0.2 + 0.4 x 0.8 = 0.44.
[0112] Weighted projection fusion: Project each view image onto the three-dimensional model to calculate the position of each pixel on the model surface. Projection transformation can be achieved through camera matrix and projection matrix to convert two-dimensional image coordinates to three-dimensional model coordinates. Use the weighted projection fusion algorithm to fuse the color and texture information of multi-view images onto the three-dimensional model. For each point on the model surface, according to its projection position in different view images, obtain the corresponding pixel value, and perform weighted fusion according to the weight. For example, for a point on the model surface, it has pixel value A in view A image with weight WA; pixel value B in view B image with weight WB. The fused pixel value is (A x WA + B x WB) / (WA + WB). Optimize the fused texture, including removing noise, smoothing transition, etc. Methods such as median filtering, Gaussian filtering, etc. can be used to remove noise in the texture, and methods such as bilinear interpolation, cubic spline interpolation, etc. can be used to smooth the texture transition, improving the overall effect of the texture.
[0113] S5, extract semantic segmentation results from the indoor-outdoor integrated surface mesh model with fine texture, combine boundary features, material properties and topological connection relationships to generate a structured surface mesh model containing wall, floor and obstacle classification.
[0114] For example, an indoor-outdoor integrated surface mesh model with fine texture has been constructed. To further improve the structured degree and semantic information of the model, it is necessary to extract semantic segmentation results from the model and combine boundary features, material properties and topological connection relationships to generate a structured surface mesh model.
[0115] Semantic Segmentation: Select a deep learning model suitable for semantic segmentation of three-dimensional models, such as 3D U-Net, PointNet++, or a model based on graph convolutional networks (GCN). These models can process three-dimensional data and perform semantic segmentation. Collect a labeled three-dimensional model dataset and train the selected deep learning model. The labeled data should include semantic categories such as walls, floors, obstacles, etc. Input the indoor-outdoor integrated surface mesh model with fine textures into the trained deep learning model to obtain the semantic labels for each vertex or voxel on the model surface.
[0116] Boundary Feature Extraction: Extract boundary features of the model surface using boundary detection algorithms (such as curvature-based, normal vector-based, or geometric feature detection based on boundaries). For example, identify the junctions of walls and floors, the boundaries of doors and windows with walls, the boundaries of stair railings and steps, etc. Optimize the extracted boundary features, including removing noisy boundaries, smoothing boundary lines, etc. Filtering algorithms or curve fitting methods can be used to make the boundaries clearer and more accurate.
[0117] Material Property Analysis: Based on the texture and color information of the model surface, classify the materials. Common materials include brick walls, glass, concrete, wood, etc. Map the classified material information to the model surface, giving each semantic region the corresponding material properties. For example, map the brick wall material to the exterior wall area, and map the glass material to the window area, etc.
[0118] Topology Connection Relationship Analysis: Construct the topology structure of the model to represent the connection relationships between different semantic regions. For example, the connection between walls, the connection between floors and walls, the connection between stairs and floors, etc. Check and optimize the topology connection relationships to ensure the rationality and completeness of the model structure. For example, repair the cracks between walls, ensure the correct connection between stairs and floors, etc.
[0119] Structured Model Generation: Integrate the semantic segmentation results, boundary features, material properties, and topology connection relationships into a structured model. Ensure that each semantic region has clear boundaries, material, and topology connection information. Output the generated structured surface mesh model in common three-dimensional model formats such as OBJ, FBX, or STL for use in different applications.
[0120] Embodiment Two
[0121] See Figure 2 A multi-modal data fusion indoor-outdoor fine three-dimensional modeling system based on the method of embodiment one, comprising:
[0122] An outdoor collection module is configured to perform unmanned aerial vehicle close-range photography, airborne laser scanning and ground supplementary photography on the building in a low-altitude area to obtain a first photo data set, a first laser point cloud data set and a POS data set of the building outdoors;
[0123] An indoor collection module is configured to continuously collect a second photo data set and a second laser point cloud data set of the building indoors by means of a laser SLAM mobile scanning system and a panoramic camera according to a pre-planned route at an indoor entrance.
[0124] A construction module is configured to perform aerial triangulation adjustment of the multi-source data set under the assistance of the POS data set and automatically construct a fine surface three-dimensional model of the indoor and outdoor integration.
[0125] In the above technical solution, in order to better use the method of one of the embodiments, the present application proposes a multi-modal data fusion indoor and outdoor fine three-dimensional modeling system, each module corresponds to each step of the above method, and the specific principle has been described in the foregoing, which will not be described here.
[0126] Embodiment three
[0127] A multi-modal data fusion indoor and outdoor fine three-dimensional modeling device comprises:
[0128] At least one processor and a memory in communication connection with the at least one processor;
[0129] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method of one of the embodiments.
[0130] In the above technical solution, in order to better run and process the method of one of the embodiments, the above method is stored in the memory, and the processor is used to execute the stored method. It should be noted that the principle and effect of each step have been described in the foregoing, which will not be expanded here.
[0131] Embodiment four
[0132] A computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method of one of the embodiments.
[0133] In the above technical solution, in order to better run and use the method of one of the embodiments, the above method is stored in the computer readable storage medium, and the processor is used to implement the above method. It should be noted that the principle and effect of each step have been described in the foregoing, which will not be expanded here.
[0134] The above merely describes some embodiments of the present application, and is not intended to limit the protection scope of the present application, and any equivalent device or equivalent process transformation, or direct or indirect application in other related technical fields, which are made by using the content of the present application specification and drawings, are also included in the patent protection scope of the present application.
Claims
1. A method for indoor and outdoor fine 3D modeling based on multimodal data fusion, characterized in that: The method comprises: Perform drone close-up photography, airborne laser scanning, and ground-based supplementary photography of buildings in low-altitude areas to obtain the first outdoor photo dataset, the first laser point cloud dataset, and the POS dataset. At the indoor entrance, a laser SLAM mobile scanning system and a panoramic camera are used to continuously collect data along a pre-planned route to obtain a second photo dataset and a second laser point cloud dataset for the building interior; With the assistance of POS datasets, aerial triangulation adjustment and solution of multi-source datasets are carried out, and a fine surface 3D model integrating indoor and outdoor elements is automatically constructed.
2. The method for indoor and outdoor fine 3D modeling based on multimodal data fusion according to claim 1, characterized in that: Using drone close-up photography, airborne laser scanning, and ground-based supplementary photography in low-altitude areas, we obtain the first photo dataset, the first laser point cloud dataset, and the POS dataset for the building's exterior. This dataset also includes: If there is a temporal and spatial inconsistency between the first photo dataset and the first laser point cloud dataset, an iterative closest point algorithm is used to perform spatiotemporal registration between them. Data synchronization is achieved through a timestamp sequence, and the spatial position deviation is corrected by combining pre-calibrated spatial coordinate parameters to obtain an outdoor registration dataset that is consistent in time and space. A third laser point cloud dataset is extracted from the outdoor registration dataset, and the point cloud density is calculated. If it is lower than a preset threshold, the missing data is inferred by combining the geometric curvature and spatial distribution characteristics of the neighboring point clouds through a generative adversarial network to obtain a density-completed laser point cloud dataset.
3. The method for indoor and outdoor fine 3D modeling based on multimodal data fusion according to claim 1, characterized in that: At the indoor entrance, the laser SLAM mobile scanning system and the panoramic camera continuously collect data along a pre-planned route to obtain a second photo dataset and a second laser point cloud dataset for the building interior, which also includes: The boundary features of walls and obstacles are extracted from the initial indoor point cloud dataset, and the point cloud integrity is calculated. If it is lower than the preset threshold, the SLAM pose graph collected by the laser SLAM mobile scanning system is used to correct the point cloud coordinate deviation to obtain an indoor point cloud dataset with optimized integrity.
4. The method for indoor and outdoor fine 3D modeling based on multimodal data fusion according to claim 1, characterized in that: With the help of POS datasets, we conduct aerial triangulation adjustment of multi-source datasets and automatically construct a fine 3D surface model integrating indoor and outdoor features, including: For the first and second laser point cloud datasets, the color and texture features of the visible light images in the first and second photo datasets were extracted. After aerial triangulation adjustment, the least squares optimization algorithm was used to perform geometric alignment with the goal of minimizing the spatial error between the point cloud coordinates and texture features, obtaining a fine surface three-dimensional model of indoor and outdoor integration with texture mapping.
5. The method for indoor and outdoor fine 3D modeling based on multimodal data fusion according to claim 1, characterized in that: Assisted by POS datasets, aerial triangulation adjustment and calculation of multi-source datasets are performed, and a fine 3D surface model integrating indoor and outdoor elements is automatically constructed. This also includes: During the aerial triangulation adjustment process, the aerial triangulation matching results are evaluated based on the depth map.
6. The method for indoor and outdoor fine 3D modeling based on multimodal data fusion according to claim 1, characterized in that: With the help of POS datasets, we conduct aerial triangulation adjustment of multi-source datasets and automatically construct a fine 3D surface model integrating indoor and outdoor features. This includes: For the indoor-outdoor integrated surface mesh model, a weighted projection fusion algorithm of multi-view images is used to perform texture mapping with view coverage and illumination consistency as weights to generate an indoor-outdoor integrated surface mesh model with fine texture; Semantic segmentation results are extracted from the indoor-outdoor integrated surface mesh model with fine textures, and a structured surface mesh model including wall, ground and obstacle classification is generated by combining boundary features, material properties and topological connection relationships.
7. A multimodal data fusion indoor and outdoor fine 3D modeling system, characterized by: Based on the method according to any one of claims 1 to 8, comprising: The outdoor acquisition module is used to perform drone close-up photography, airborne laser scanning, and ground supplementary photography of buildings in low-altitude areas to obtain the first photo dataset, the first laser point cloud dataset, and the POS dataset outdoors. The indoor acquisition module is used to continuously acquire data along a pre-planned route at the indoor entrance using a laser SLAM mobile scanning system and a panoramic camera to obtain a second photo dataset and a second laser point cloud dataset for the interior of the building; The construction module is used to carry out aerial triangulation adjustment of multi-source datasets with the assistance of POS datasets, and automatically construct a fine surface 3D model that integrates indoor and outdoor elements.
8. A multimodal data fusion indoor and outdoor fine three-dimensional modeling device, characterized by: include: at least one processor and a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Semantic three-dimensional model construction method and system of building, terminal and storage medium
CN121120952A
Artificial intelligence AI auxiliary space checking structure extraction method and system
CN121686225A