Multi-machine collaborative three-dimensional map updating method based on mobile augmented reality
By employing a multi-machine collaborative 3D map update method, and utilizing AR glasses and deep neural networks for high-precision 3D modeling, the current availability and accuracy issues of real-world 3D model updates in existing technologies are resolved. This enables efficient and automated 3D map updates and accurate virtual-real registration, meeting the high-frequency update needs of smart cities and digital twins.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-11
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies suffer from problems such as poor timeliness, low accuracy, low efficiency, and inconsistent data fusion in the dynamic updating of real-world 3D models and the application of spatial information fusion. These technologies are insufficient to meet the high-frequency and high-timeliness update requirements of fields such as smart cities and digital twins.
A multi-machine collaborative 3D map update method based on mobile augmented reality is adopted. Data is collected by integrating binocular cameras, GPS modules and IMU modules through multiple AR glasses. Combined with deep neural networks and multi-view stereo geometry algorithms, high-precision 3D modeling and data fusion are achieved, a unified spatiotemporal benchmark for GIS and AR is constructed, and virtual and real registration is carried out.
It enables efficient and automated 3D map updates, reduces hardware costs and manual intervention, improves modeling accuracy and stability, ensures precise matching of virtual and real spaces, and meets the high-frequency update needs of smart cities and digital twins.
Smart Images

Figure CN121837528A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of three-dimensional reconstruction, in particular to a multi-machine cooperative three-dimensional map updating method based on mobile augmented reality. BACKGROUND
[0002] With the promotion of digital economy and smart city construction, from real three-dimensional to digital twin has become a core demand. Three-dimensional models, as the basic data of location services, are the key to supporting digital twin applications. Their timeliness directly affects the sustainable development of spatial location services. Currently, a large number of static three-dimensional models have been widely used in various industries. At the same time, the rapid development of intelligent devices has made applications based on visual augmented reality (AR) stand out. This technology has the advantages of low cost, high efficiency, easy understanding, and strong timeliness, providing a new direction for dynamic updating of three-dimensional models.
[0003] However, the existing technology still has many problems and defects in the field of real three-dimensional model dynamic updating and spatial information fusion application. Among them, static three-dimensional models generally have poor timeliness, that is, they cannot reflect the dynamic changes of the real scene in real time, making it difficult to meet the needs of smart cities, digital twins, and other fields for high-frequency and high-timeliness updates of spatial data. Although AR technology has been applied in visual display and simple interaction scenarios, its application in three-dimensional model updating with surveying and mapping level precision is still immature, and there is a lack of effective fusion design with professional spatial data acquisition and modeling processes. In the fusion process of multi-source spatial data, the time and space benchmarks are not unified, which can easily lead to spatial position deviation after data fusion, affecting the precise matching effect of virtual and real spaces. Traditional three-dimensional reconstruction and data processing mostly use single-device serial processing mode, lacking automated and batch processing mechanisms. When facing massive acquisition data of large-scale scenes such as cities, the processing efficiency is extremely low, making it difficult to support rapid updating needs. At the same time, existing image feature extraction, matching, and point cloud optimization technologies are easily affected by environmental factors such as light changes, view differences, and weak textures in complex actual scenarios, resulting in poor accuracy and stability of three-dimensional modeling. The quality of the generated model is difficult to meet the requirements of professional applications. SUMMARY
[0004] The purpose of the present application is to provide a multi-machine cooperative three-dimensional map updating method based on mobile augmented reality to solve the above technical problems.
[0005] To achieve the above purpose, the present application provides a multi-machine cooperative three-dimensional map updating method based on mobile augmented reality, comprising the following steps: S1: Collecting the binocular video image sequence, GPS data and IMU data of the target area by the AR glasses with multiple integrated binocular cameras, GPS modules and IMU modules, and inputting the two-dimensional POI attribute data through voice instructions to locate the current GPS position of the device, and then obtaining the initial collection data with geographical labels; S2: Based on the initial collection data obtained in S1, the original video images collected are frame-picked and sorted, and the average latitude and longitude of each AR glasses and the physical distance between the center points of the devices are calculated according to the GPS data. The distance threshold is set to determine whether the target corresponding to the collection data is the same object, and the data grouping result is obtained and distributed to the corresponding processing unit. S3: Based on the data grouping result obtained in S2, a hybrid pairing strategy of sliding window + closed loop detection is used to generate a sparse pairing list. The image feature points and descriptors are extracted by the deep neural network model integrated in the HLOC hierarchical positioning framework, and then the feature matching is performed by the attention mechanism network and the outliers are removed. Combined with the incremental SfM solution, the sparse point cloud data is obtained. S4: Based on the camera distortion parameters synchronously output by the incremental SfM in S3 when solving the sparse point cloud, the frame-picked images after grouping in S2 are de-distorted, and then the depth map is estimated and fused into the initial dense point cloud by the multi-view stereo geometry algorithm. The initial dense point cloud is optimized by clustering algorithm and statistical filtering to obtain the pure dense point cloud data. S5: Based on the pure dense point cloud data obtained in S4, a triangular mesh model is generated by Poisson surface reconstruction, and the view selection is optimized by Markov random field to complete texture mapping, and a high-precision three-dimensional model is obtained. S6: Based on the high-precision three-dimensional model obtained in S5 and the two-dimensional POI attribute data obtained in S1, a unified space-time reference conversion matrix of GIS geographical coordinate system and AR local tracking coordinate system is constructed, and virtual-real registration is realized by sequential conversion of multiple coordinate systems. The updated three-dimensional model and POI attribute data are pushed to the AR glasses and the front-end visualization platform to obtain real-time AR visualization navigation and map update results.
[0006] Preferably, in S1, the association of two-dimensional POI attribute data and GPS data is automatically completed by the spatial information augmented reality visualization APP built-in the AR glasses.
[0007] Preferably, in S2, the distance threshold is adjusted adaptively according to the scene scale of the target area, and the judgment rule is: if the physical distance between the center points of the devices does not exceed the threshold, it is determined as the same building, and the images after frame-picking of multiple AR glasses are put into the same folder for separate modeling; otherwise, it is determined as not the same building, and each is modeled separately in different folders.
[0008] Preferably, the specific steps of S3 include: S31, based on the data grouping result obtained in S2, a hybrid spatio-temporal constraint pairing strategy of sliding window + closed loop detection is used, a window threshold is set to cover consecutive frames, a loop threshold is set to enhance the head-tail geometric constraint, and a sparse pairing list is generated; S32, the SuperPoint deep neural network model integrated by the HLOC hierarchical positioning framework is used to infer the frame images grouped in S2, and automatically extract image feature points and corresponding descriptors which are highly robust to illumination changes, weak textures and view transformations; S33, run the HLOC-based automated script to drive the LightGlue attention mechanism network to automatically match the feature points and descriptors extracted in S32, and automatically remove outliers through geometric verification, and output high-quality sparse feature correspondence; S34, using the high-quality sparse feature correspondence obtained in S33, incremental motion recovery structure SfM solving is performed to generate sparse point cloud data.
[0009] Preferably, the steps of the incremental motion recovery structure SfM solving in S34 are: recovering three-dimensional space points by triangulation, simultaneously solving the 6-DoF camera pose and internal parameter of each frame image by joint optimization, generating an initial sparse scene structure, and then obtaining the sparse point cloud data.
[0010] Preferably, the specific steps of S4 include: S41, based on the incremental SfM solving result corresponding to the sparse point cloud data obtained in S34, extract the estimated camera distortion parameters, resample and de-distort the grouped original images to generate image data under the standard pinhole camera model; S42, using the PatchMatchStereo algorithm, based on the standard pinhole camera model image data obtained in S41, combining the photometric consistency constraint and the geometric consistency constraint, the depth map and the normal map of the image are estimated pixel by pixel; S43, geometrically checking the depth map obtained in S42, fusing the depth map that passes the check into an initial dense point cloud, and setting a minimum pixel consistency threshold to filter outliers in the initial dense point cloud; S44, calling a self-defined Python script, performing voxel downsampling on the filtered dense point cloud obtained in S43 to reduce data redundancy, and then introducing the DBSCAN density clustering algorithm and statistical filtering to automatically segment and retain the maximum connected body of the point cloud, and remove residual outliers caused by the sky background or dynamic objects, to obtain pure dense point cloud data.
[0011] Preferably, the specific steps of S5 include: S51, based on the pure dense point cloud data obtained in S44, a Poisson equation is used to solve an indicator function to reconstruct a Poisson surface to achieve a topological conversion from a discrete point cloud to a continuous surface, and then a closed triangular mesh model is generated; S52, a Markov random field (MRF) optimization strategy is used for optimal view selection; S53, the texture information of the multi-frame images selected in S52 is mapped to the triangular mesh model generated in S51 to generate a high-resolution texture atlas and a high-precision three-dimensional model with UV coordinates.
[0012] Preferably, based on S51, a temporary isolated processing sandbox is constructed to transmit necessary sparse camera parameters only through a soft link or a copy mode, thereby ensuring data processing stability.
[0013] Preferably, in S6, virtual-real registration is achieved through multi-coordinate system conversion based on a unified space-time reference conversion matrix, and the specific conversion process is as follows: S61, the three-dimensional coordinates in the virtual object coordinate system are converted into three-dimensional coordinates in the augmented reality space coordinate system , and the conversion formula is: ; wherein, is a homogeneous transformation matrix between coordinate systems; is a 3x3 rotation matrix; is a 3x1 translation vector; S62, based on the real-time pose of the camera, the three-dimensional coordinates in the augmented reality space coordinate system are converted into three-dimensional coordinates in the user observation coordinate system , and the conversion formula is: ; S63, the three-dimensional coordinates in the user observation coordinate system are projected to the display plane coordinate system to obtain two-dimensional pixel coordinates, and augmented reality registration is achieved, and the projection formula is: ; wherein, is the homogeneous pixel coordinates of the corresponding point in the display plane coordinate system; is the depth value of the three-dimensional point in the user observation coordinate system; is the pixel scale coefficient of the camera image sensor in the horizontal direction; is the pixel scale coefficient of the camera image sensor in the vertical direction; is the principal point horizontal coordinate of the camera image plane; is the principal point vertical coordinate of the camera image plane; is the focal length of the camera. Preferably, in S6, the high-precision three-dimensional model updated in S63 is pushed to the front-end visualization platform and registered to the corresponding real geographical position to realize rapid map updating, and the updated two-dimensional POI attribute data is pushed to the AR glasses, and the AR glasses dynamically render the POI labels on the screen according to the real-time pose and field of view range and the perspective rule of near large and far small to obtain real-time AR visualization navigation and map updating results.
[0014] Therefore, the application has the beneficial effects of the above-mentioned multi-machine collaborative three-dimensional map updating method based on mobile augmented reality, which is characterized by: 1. The AR glasses are used as an integrated collection terminal integrating a binocular camera, a GPS module, an IMU module and a voice interaction function to replace traditional expensive professional surveying and mapping equipment, synchronously collect full-factor data required for three-dimensional modeling and two-dimensional POI attribute data input by voice and automatically associate geographical labels, realize deep and seamless integration of AR technology and three-dimensional modeling, reduce the hardware cost and technical threshold of three-dimensional map updating, complete data collection without the need for professional surveying and mapping personnel to operate, ensure the comprehensiveness and relevance of the collected data, lay an efficient data foundation for subsequent high-precision modeling and real-time updating, and significantly simplify the data collection process.
[0015] 2. Based on the GPS data collected by multiple AR glasses, the average latitude and longitude of the calculation equipment and the physical distance of the center point are calculated, a distance threshold is adaptively adjusted according to the scene scale, the collection target is automatically determined and the data is distributed for parallel modeling without manual intervention in data grouping, the efficiency bottleneck of traditional single-device serial processing is broken through, the processing capacity of city-level massive image data is effectively improved, the automatic and efficient operation of multi-device collaborative modeling is realized, the system can continuously and stably respond to high-frequency iterative updating tasks, the updating cycle of the three-dimensional map is greatly shortened, and the engineering practicability of the technical scheme is enhanced.
[0016] 3. By constructing a four-step coordinate system conversion matrix of virtual objects-AR space-user observation-display plane, a unified space-time reference of GIS and AR is established, a full-process data closed loop of "collection-upload-pushing" is formed, the drift and misplacement problem of GIS data and AR scene fusion in traditional technology is solved, accurate virtual-real registration is realized, the updated high-precision three-dimensional model and POI attribute data can be pushed to the front-end visualization platform and AR glasses in real time, near-large-far-small perspective rule navigation and interactive experience that fit the real scene are provided for users, the usability of spatial information services and the present situation of digital twin systems are improved, and a complete and efficient spatial information service system is constructed.
[0017] The technical solutions of the application will be further described in detail below with reference to the drawings and embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0018] Figure 1 Structure diagram of AR glasses used in the mobile augmented reality based multi-machine cooperative three-dimensional map updating method provided by the present application; Figure 2 System overall architecture diagram of the mobile augmented reality based multi-machine cooperative three-dimensional map updating method provided by the present application; Figure 3 Two-dimensional POI attribute information collection flowchart in the mobile augmented reality based multi-machine cooperative three-dimensional map updating method provided by the present application; Figure 4 Three-dimensional reconstruction flowchart in the mobile augmented reality based multi-machine cooperative three-dimensional map updating method provided by the present application.
[0019] Reference numerals 1, LOGO light; 2, black and white camera; 3, universal key; 4, speaker hole one; 5, color camera; 6, power key; 7, speaker hole two. DETAILED DESCRIPTION
[0020] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the embodiments of the present application are further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the embodiments of the present application and are not used to limit the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of the present application. Examples of embodiments are shown in the drawings, wherein the same or similar reference numerals represent the same or similar elements or elements having the same or similar functions throughout.
[0021] It should be noted that the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or server comprising a series of S or units does not have to be limited to those clearly listed S or units, but can include other S or units not clearly listed or inherent to these processes, methods, products or devices.
[0022] The embodiments of the present application are described in detail below in combination with the drawings.
[0023] There are many shortcomings in the prior art in real scene three-dimensional model updating and spatial information application. The static three-dimensional model is poor in present situation and cannot adapt to high frequency dynamic updating demand. The application of AR technology in three-dimensional model updating in the field of surveying and mapping level precision is not mature and is not effectively combined with professional spatial data processing flow. The time and space reference of multi-source spatial data fusion is not unified, which easily causes spatial position deviation and affects virtual-real matching effect. Three-dimensional reconstruction mostly adopts single device serial processing mode, lacks automatic batch processing mechanism and has low processing efficiency when facing large-scale scene mass data. At the same time, the image feature extraction and matching technology lacks robustness in complex scenes and is easily disturbed by environmental factors, resulting in poor precision and stability of three-dimensional modeling, which is difficult to meet the professional application requirements in the fields of smart city and digital twin.
[0024] The present application is designed based on the above analysis, and the AR glasses used in the present application are shown in Figures 1-4 The mobile augmented reality-based multi-machine cooperative three-dimensional map updating method comprises the following steps: S1: Collecting the binocular video image sequence, GPS data and IMU data of the target area through the AR glasses integrated with double binocular cameras, GPS modules and IMU modules, inputting the two-dimensional POI attribute data through voice instructions to locate the current GPS position of the device, obtaining the initial collection data with geographical labels, and uploading the initial collection data to the back end in real time; Figure 3 In S1, the association of the two-dimensional POI attribute data and the GPS data is automatically completed through the spatial information augmented reality visualization APP built in the AR glasses.
[0025] Specifically, the AR glasses used in the present application are shown in Figure 1 The specific structure of the AR glasses used in the present application comprises: a LOGO lamp 1 is arranged on the upper part of the right temple, a black and white camera 2 for collecting grayscale visual data is arranged on the upper part of the right frame, a universal key 3 for main function interaction control of the device and a first loudspeaker hole 4 for audio output are arranged on the middle and lower parts of the right temple respectively; a color camera 5 for collecting color visual data is arranged on the outer side of the left frame to realize binocular visual collection, a power key 6 and a second loudspeaker hole 7 are arranged on the middle and lower parts of the left temple respectively, and audio output is realized through the symmetrically arranged loudspeaker holes.
[0026] S2: Based on the GPS data in the initial collection data obtained in S1, the original video images collected are frame-picked and sorted, the average longitude and latitude of each AR glasses and the physical distance between the center points of the devices are calculated according to the GPS data, whether the targets corresponding to the collection data are the same object is determined by setting a distance threshold, the data grouping result is obtained and distributed to the corresponding processing unit; In S2, the distance threshold is adaptively adjusted according to the scene scale of the target area, and the judgment rule is shown in Figure 4 If the center point physical distance between the devices does not exceed the threshold value, it is determined that they are in the same building, and the images of the multiple AR glasses after frame extraction are placed in the same folder for separate modeling. Otherwise, it is determined that they are not in the same building, and each is modeled separately in different folders.
[0027] S3: Based on the data grouping result obtained in S2, a hybrid pairing strategy of sliding window + closed loop detection is used to generate a sparse pairing list. Image feature points and descriptors are extracted by a deep neural network model integrated by the HLOC hierarchical positioning framework, and then feature matching and outlier elimination are performed by an attention mechanism network. Combined with incremental SfM solving, sparse point cloud data is obtained. The specific steps of S3 include: S31, based on the data grouping result obtained in S2, a hybrid spatio-temporal constraint pairing strategy of sliding window + closed loop detection is used. The window threshold is set to cover consecutive frames, and the loop threshold is set to enhance the head-tail geometric constraint. The calculation redundancy of full connection matching is abandoned, and a sparse pairing list is generated. The window threshold is Window=20, and the loop threshold is Loop=40. S32, the grouped frame-extracted images of S2 are inferred by the SuperPoint deep neural network model integrated by the HLOC hierarchical positioning framework. There is no need to pre-process the images for distortion correction. Image feature points and corresponding descriptors with high robustness to illumination changes, weak textures and perspective transformations are automatically extracted. The HLOC hierarchical positioning framework is an open-source hierarchical positioning framework for visual positioning and three-dimensional reconstruction. It provides a standardized integrated environment, algorithm scheduling and end-to-end process management capability for various visual feature extraction and matching algorithms, and can efficiently adapt to different deep neural network models and simplify the engineering application link. The SuperPoint deep neural network model is a visual feature extraction model based on convolutional neural network, designed for automatic generation of image feature points and corresponding descriptors. By using the HLOC hierarchical positioning framework as the algorithm integration and running carrier, the grouped frame-extracted images are directly input. The SuperPoint deep neural network model uses multi-layer convolutional neural network to perform hierarchical convolution, feature mapping and adaptive filtering on image pixels. Without pre-distortion correction of the image, it can automatically detect key feature points in the image with strong robustness to illumination changes, weak textures and perspective transformations, and generate unique high-identification descriptors for each feature point.
[0028] S33, running the HLOC-based automation script, driving the LightGlue attention mechanism network to script the matching of the feature points and descriptors extracted by S32, automatically traversing the binocular image sequence, synchronously calculating the stereo matching relationship between the left and right images and the time sequence matching relationship between adjacent frames, automatically eliminating outliers through geometric verification, and outputting high-quality sparse feature correspondence; The LightGlue attention mechanism network is a lightweight deep feature matching network equipped with an attention mechanism, which can efficiently adapt to visual features extracted by models such as SuperPoint. It is the core algorithm module for high-precision feature matching in the HLOC hierarchical positioning framework. The LightGlue attention mechanism network dynamically models the correlation between feature points and descriptors in different images based on the attention mechanism, autonomously focuses on high-confidence feature correlations and suppresses noise interference. Under the scheduling of the HLOC automation script, it performs global traversal on the binocular image sequence, synchronously completes stereo matching of left and right images and time sequence matching of adjacent frame images, and combines geometric verification rules to check the preliminary matching results. It automatically eliminates mismatched outliers caused by sudden changes in light, repeated textures, and perspective distortion, and finally outputs high-quality sparse feature correspondence that meets the accuracy requirements of three-dimensional reconstruction.
[0029] S34, using the high-quality sparse feature correspondence obtained by S33, performing incremental motion recovery structure SfM solving to generate sparse point cloud data.
[0030] The steps of S3 are: recovering three-dimensional space points through triangulation, simultaneously optimizing the 6-DoF camera pose and intrinsic parameters of each frame of image to generate an initial sparse scene structure, and then obtaining sparse point cloud data.
[0031] S4: Based on the camera distortion parameters synchronously output by the incremental SfM of S3 when solving sparse point cloud, the frame-drawing images after grouping in S2 are de-distorted, and then the depth map is estimated and fused into an initial dense point cloud through multi-view stereo geometry algorithm. The initial dense point cloud is optimized by clustering algorithm and statistical filtering to obtain pure dense point cloud data. The specific steps of S4 include: S41, based on the incremental SfM solving result corresponding to the sparse point cloud data obtained by S34, extracting the estimated camera distortion parameters, resampling and de-distorting the grouped original images to generate image data under the standard pinhole camera model, eliminating perspective errors to adapt to subsequent multi-view stereo geometry calculation; S42, based on the standard pinhole camera model image data obtained in S41, a depth map and a normal map of the image are estimated pixel by pixel by using a PatchMatchStereo algorithm combined with photometric consistency constraint and geometric consistency constraint; the PatchMatchStereo algorithm is a high-efficiency multi-view stereo matching dense reconstruction algorithm, which can quickly generate high-precision depth information and surface normal vector from binocular or multi-view images, and adapt to the corrected images of the standard pinhole camera model; taking an image local pixel block as a matching basic unit, through a standardized process of random initialization of disparity plane hypothesis, space propagation, view propagation and iterative optimization, the color and gray similarity of corresponding pixel blocks under different views are checked by combining with the photometric consistency constraint, and the spatial smoothness and disparity continuity of the depth field are guaranteed by relying on the geometric consistency constraint, the matching result is iteratively optimized pixel by pixel, the calculated disparity value is converted into a depth value combined with the camera intrinsic parameter, and a three-dimensional surface normal vector corresponding to each pixel is solved synchronously, and finally a dense and robust depth map and normal map are output; S43, the depth map obtained in S42 is geometrically checked, the checked depth map is fused into an initial dense point cloud, and a minimum pixel consistency threshold is set to filter out outliers in the initial dense point cloud; S44, a self-defined Python script is called to perform voxel downsampling processing on the filtered dense point cloud obtained in S43 to reduce data redundancy, and then a DBSCAN density clustering algorithm and statistical filtering are introduced to automatically segment and retain the maximum connected body of the point cloud and to remove residual outliers caused by the sky background or dynamic objects, so as to obtain pure dense point cloud data.
[0032] S5: based on the pure dense point cloud data obtained in S4, a triangular mesh model is generated by Poisson surface reconstruction, and Markov random field optimization view selection is used to complete texture mapping, so as to obtain a high-precision three-dimensional model; The specific steps of S5 include: S51, based on the pure dense point cloud data obtained in S44, an indicator function is solved by using a Poisson equation to reconstruct a Poisson surface, so as to realize the topological conversion from a discrete point cloud to a continuous surface, and then a closed triangular mesh model is generated; the indicator function is solved by using the Poisson equation to reconstruct the Poisson surface, so as to realize the generation of a continuous and smooth closed surface from a discrete dense point cloud, so as to adapt to the topological reconstruction requirement of the pure dense point cloud; based on the pure dense point cloud and its normal vector field obtained in S44, a Poisson equation with the indicator function as an unknown quantity is constructed by spatial integration of the normal vector field, the continuous and smooth indicator function surface that fits the spatial distribution and geometric characteristics of the point cloud is obtained by solving the partial differential equation, the topological conversion from the discrete point set to the continuous surface is completed, and finally the continuous surface with complete topological structure and smooth surface is generated by isosurface extraction and triangulation division of the continuous indicator function surface.
[0033] Based on S51, a temporary isolated processing sandbox is constructed to transmit necessary sparse camera parameters only through soft links or copying, and to physically shield the implicit occupation of the memory by the depth map, thereby ensuring the stability of data processing.
[0034] S52, the Markov Random Field (MRF) optimization strategy is used for optimal view selection to solve the problems of texture joint misalignment and color inconsistency that may occur in multi-view projection; wherein the Markov Random Field (MRF) optimization strategy is a graph model energy optimization method constructed based on the Markov Random Field theory, which converts the view selection problem of three-dimensional grid cells into a global energy minimization problem, takes the grid cells or projection pixels as the nodes of MRF, and the view selection correlation between adjacent nodes as the edges, constructs an energy function containing data items and smoothing items, wherein the data items are used to measure the texture projection matching degree and color fidelity of a single node in the candidate view, to ensure the projection adaptability of the view and the cell, and the smoothing items are used to constrain the view selection of adjacent nodes to maintain spatial continuity and avoid the abruptness of local view switching, the minimum value of the energy function is solved by an iterative optimization algorithm, and the globally optimal view allocation is realized, thereby unifying the view selection logic of multi-view projection, eliminating the misalignment problem at the texture splicing place, weakening the color difference between different views, and ensuring the visual continuity and overall consistency of the three-dimensional model surface texture mapping.
[0035] S53, the texture information of the multi-frame images selected and optimized in S52 is mapped to the triangular mesh model generated in S51 to generate a high-resolution texture atlas and a high-precision three-dimensional model with UV coordinates.
[0036] S6: based on the high-precision three-dimensional model obtained in S5 and the two-dimensional POI attribute data obtained in S1, a unified space-time reference conversion matrix of the GIS geographic coordinate system and the AR local tracking coordinate system is constructed, virtual-real registration is realized through sequential conversion of multiple coordinate systems, and the updated three-dimensional model and POI attribute data are pushed to the AR glasses and the front-end visualization platform to obtain real-time AR visualization navigation and map update results.
[0037] In S6, based on the unified space-time reference conversion matrix, virtual-real registration is realized through sequential conversion of multiple coordinate systems, and the specific conversion process is as follows: S61, the three-dimensional coordinates in the virtual object coordinate system are converted into the three-dimensional coordinates in the augmented reality space coordinate system , and the conversion formula is: ; wherein, is a homogeneous transformation matrix between coordinate systems, used to realize coordinate conversion from the original coordinate system to the augmented reality space coordinate system, and is a 4x4 matrix containing rotation and translation information; is a 3x3 rotation matrix, used to represent the rotation relationship between the virtual object coordinate system and the augmented reality space coordinate system, and to ensure the length and angle in the coordinate transformation unchanged; is a 3x1 translation vector, representing the translation amount of the virtual object coordinate system origin relative to the augmented reality space coordinate system origin, corresponding to the displacement in X, Y, and Z axis directions; S62, based on the real-time pose of the camera, converts the three-dimensional coordinates in the augmented reality space coordinate system to three-dimensional coordinates in the user observation coordinate system , and the conversion formula is: ; S63, projects the three-dimensional coordinates in the user observation coordinate system to the display plane coordinate system to obtain two-dimensional pixel coordinates, realizes augmented reality registration, and the projection formula is: ; wherein, is the homogeneous pixel coordinates of the corresponding point in the display plane coordinate system; is the depth value of the three-dimensional point in the user observation coordinate system, i.e. the vertical distance from the point to the camera image plane; is the pixel scale coefficient of the camera image sensor in the horizontal direction; is the pixel scale coefficient of the camera image sensor in the vertical direction; is the principal point horizontal coordinate of the camera image plane; is the principal point vertical coordinate of the camera image plane; is the focal length of the camera.
[0038] The updated high-precision three-dimensional model in S63 is pushed to the front-end visualization platform and registered to the corresponding real geographical location to realize fast map updating; at the same time, the updated two-dimensional POI attribute data is pushed to the AR glasses, and the AR glasses dynamically render the POI label on the screen according to the real-time pose and field of view range, the perspective rule of near large and far small, provide AR visualization navigation function, form a full-process data closed loop of collection- upload-pushing, to obtain real-time AR visualization navigation and map updating results.
[0039] In summary, the application reduces the hardware threshold and labor cost of three-dimensional model acquisition and updating, improves the efficiency of multi-device collaborative processing, and the precision and stability of modeling in complex scenes, realizes the rapid updating of three-dimensional models and the accurate fusion of AR visual navigation, effectively guarantees the timeliness of spatial data and the accuracy of virtual-real interaction, forms an efficient, automated, and highly adaptive real scene three-dimensional model updating and AR navigation fusion scheme, and can fully meet the core needs of high-frequency updating, high-precision modeling and accurate spatial information services in the field of smart city, digital twin and the like.
[0040] Finally, it should be noted that: the above examples are only used to illustrate the technical solutions of the present application, but not to limit it, although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that: the technical solutions of the present application can still be modified or replaced by the equivalent, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present application.
Claims
1. A multi-machine collaborative 3D map update method based on mobile augmented reality, characterized in that: Includes the following steps: S1: Using multiple AR glasses that integrate binocular cameras, GPS modules, and IMU modules, binocular video image sequences, GPS data, and IMU data of the target area are collected respectively. Two-dimensional POI attribute data is input through voice commands to locate the current GPS position of the device, thereby obtaining the initial collected data with geographic tags. S2: Based on the GPS data in the initial acquisition data obtained in S1, the original video images are extracted and sorted, and the average latitude and longitude of each AR glasses and the physical distance between the center points of the devices are calculated according to the GPS data. By setting a distance threshold, it is determined whether the target corresponding to the acquisition data is the same object, and the data grouping results are obtained and distributed to the corresponding processing units. S3: Based on the data grouping results obtained in S2, a sparse pairing list is generated using a hybrid pairing strategy of sliding window + loop closure detection. Image feature points and descriptors are extracted through a deep neural network model integrated by the HLOC hierarchical localization framework. Then, feature matching and outlier removal are performed through an attention mechanism network. Combined with incremental SfM solution, sparse point cloud data is obtained. S4: The incremental SfM based on S3 synchronously outputs camera distortion parameters when solving sparse point clouds. It performs distortion removal processing on the frame-sampling images after grouping in S2, and then estimates the depth map of the distortion-removed images through a multi-view stereo geometry algorithm and fuses them into an initial dense point cloud. Clustering algorithm and statistical filtering are used to optimize the initial dense point cloud to obtain clean dense point cloud data. S5: Based on the clean and dense point cloud data obtained in S4, a triangular mesh model is generated by reconstructing a Poisson surface, and the view selection is optimized by using a Markov random field to complete the texture mapping and obtain a high-precision 3D model. S6: Based on the high-precision 3D model obtained in S5 and the 2D POI attribute data obtained in S1, a unified spatiotemporal reference transformation matrix between the GIS geographic coordinate system and the AR local tracking coordinate system is constructed. Virtual and real registration is achieved through sequential transformation of multiple coordinate systems. The updated 3D model and POI attribute data are pushed to AR glasses and the front-end visualization platform to obtain real-time AR visualization navigation and map update results.
2. The multi-machine collaborative 3D map update method based on mobile augmented reality according to claim 1, characterized in that: In S1, the association between 2D POI attribute data and GPS data is automatically completed through the spatial information augmented reality visualization APP built into the AR glasses.
3. The multi-machine collaborative 3D map update method based on mobile augmented reality according to claim 2, characterized in that: In S2, the distance threshold is adaptively adjusted according to the scene scale of the target area, and the judgment rule is: if the physical distance between the center points of the devices does not exceed the threshold, they are judged to be the same building, and the images extracted from multiple AR glasses are put into the same folder for separate modeling. Conversely, if they are not the same building, they will be modeled separately in different folders.
4. The multi-machine collaborative 3D map update method based on mobile augmented reality according to claim 3, characterized in that: The specific steps of S3 include: S31. Based on the data grouping results obtained in S2, a hybrid spatiotemporal constraint pairing strategy of sliding window + loop closure detection is adopted. A window threshold is set to cover continuous frames, and a loop closure threshold is set to enhance the geometric constraints at the beginning and end, generating a sparse pairing list. S32. The SuperPoint deep neural network model integrated by the HLOC hierarchical localization framework is used to infer the grouped frame images of S2, and automatically extract image feature points and corresponding descriptors that are highly robust to changes in illumination, weak texture and viewpoint. S33. Run an automated script based on HLOC to drive the LightGlue attention mechanism network to perform script-based fully automatic matching of feature points and descriptors extracted in S32, and automatically remove outliers through geometric verification to output high-quality sparse feature correspondences. S34. Using the high-quality sparse feature correspondence obtained in S33, perform incremental motion recovery structure (SfM) calculation to generate sparse point cloud data.
5. The multi-machine collaborative 3D map update method based on mobile augmented reality according to claim 4, characterized in that: The incremental motion recovery structure (SfM) solution in S34 is as follows: three-dimensional spatial points are recovered through triangulation, and the 6-DoF camera pose and intrinsic parameters of each frame of image are jointly optimized to generate an initial sparse scene structure, thereby obtaining sparse point cloud data.
6. The multi-machine collaborative 3D map update method based on mobile augmented reality according to claim 5, characterized in that: The specific steps of S4 include: S41. Based on the incremental SfM solution results corresponding to the sparse point cloud data obtained in S34, extract the estimated camera distortion parameters, resample and dedistort the grouped original images, and generate image data under the standard pinhole camera model. S42. Using the PatchMatchStereo algorithm, based on the standard pinhole camera model image data obtained in S41, and combining photometric consistency constraints and geometric consistency constraints, the depth map and normal map of the image are estimated pixel by pixel. S43. Perform geometric verification on the depth map obtained in S42, fuse the verified depth maps into an initial dense point cloud, and set a minimum pixel consistency threshold to filter out outliers in the initial dense point cloud. S44. Call a custom Python script to perform voxel downsampling on the filtered dense point cloud obtained in S43 to reduce data redundancy. Then, introduce the DBSCAN density clustering algorithm and statistical filtering to automatically segment and retain the largest connected entity of the point cloud and remove residual outliers caused by the sky background or dynamic objects to obtain clean dense point cloud data.
7. The multi-machine collaborative 3D map update method based on mobile augmented reality according to claim 6, characterized in that: The specific steps of S5 include: S51. Based on the clean and dense point cloud data obtained in S44, the indicator function is solved using the Poisson equation to reconstruct the Poisson surface, so as to realize the topological transformation from discrete point cloud to continuous surface, and then generate a closed triangular mesh model. S52. The best view is selected by using the Markov Random Field (MRF) optimization strategy. S53. Map the texture information of the multi-frame images optimized and selected in S52 onto the triangular mesh model generated in S51 to generate a high-resolution texture atlas and a high-precision 3D model with UV coordinates.
8. The multi-machine collaborative 3D map update method based on mobile augmented reality according to claim 7, characterized in that: Based on S51, a temporary isolated processing sandbox is built to transmit necessary sparse camera parameters only through soft links or copies, thereby ensuring the stability of data processing.
9. The multi-machine collaborative 3D map update method based on mobile augmented reality according to claim 8, characterized in that: In S6, virtual-real registration is achieved through sequential transformation across multiple coordinate systems based on a unified spatiotemporal reference transformation matrix. The specific transformation process is as follows: S61. Transfer the three-dimensional coordinates in the virtual object coordinate system Convert to 3D coordinates in augmented reality space coordinate system The conversion formula is: ; in, This is the homogeneous transformation matrix between coordinate systems; It is a 3×3 rotation matrix; It is a 3×1 translation vector; S62. Based on the real-time pose of the camera, the three-dimensional coordinates in the augmented reality coordinate system will be... Convert to three-dimensional coordinates in the user's viewing coordinate system The conversion formula is: ; S63. The three-dimensional coordinates in the user's observation coordinate system Projecting onto the display plane coordinate system yields two-dimensional pixel coordinates, enabling augmented reality registration. The projection formula is: ; in, To display the homogeneous pixel coordinates of the corresponding point in the planar coordinate system; To allow users to observe the depth values of 3D points in the coordinate system; is the pixel scale factor of the camera image sensor in the horizontal direction; is the pixel scale factor of the camera image sensor in the vertical direction; The horizontal coordinates of the principal point on the camera image plane; The vertical coordinates of the principal point on the camera image plane; This refers to the camera's focal length.
10. The multi-machine collaborative 3D map update method based on mobile augmented reality according to claim 9, characterized in that: In S6, the updated high-precision 3D model from S63 is pushed to the front-end visualization platform and registered to the corresponding real geographical location to achieve rapid map updates. At the same time, the updated 2D POI attribute data is pushed to the AR glasses. The AR glasses dynamically render POI labels on the screen according to the perspective rule of near objects appearing larger and far objects appearing smaller, based on the real-time pose and field of view, to obtain real-time AR visualization navigation and map update results.
Citation Information
Cited By
Incremental SfM real-time point cloud reconstruction method for single-region unmanned aerial vehicle cluster
CN122115752A
Single-region unmanned aerial vehicle cluster incremental SfM real-time point cloud reconstruction method
CN122115752B