Large-scale three-dimensional scene reconstruction method, system, equipment and medium
By synchronizing the positioning device and the image acquisition device in time, geospatial coordinates and video frame images are obtained. Key image frames are extracted based on the spatial distance, density features and global features between video frames. Bundle adjustment is used to reconstruct 3D point clouds and optimize image pose. This solves the problems of high equipment cost, insufficient reconstruction accuracy and efficiency in the existing technology, and realizes large-scale 3D scene reconstruction with low cost and high accuracy.
Patent Information
- Application Number
- CN202511015007.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-11-21
AI Technical Summary
Existing 3D reconstruction technologies suffer from high equipment costs, insufficient reconstruction accuracy and efficiency in consumer-grade scenarios. They are particularly inefficient and time-consuming when processing large-scale data, and are prone to reconstruction failures when dealing with areas with weak textures.
By synchronizing the positioning device and the image acquisition device in time, geospatial coordinates and video frame images are obtained. Key image frames are extracted based on the spatial distance between video frames, dense features and global features. Bundle adjustment is used to reconstruct 3D point clouds and optimize image pose. In weak texture areas, an image matching network without keypoint detection is used to estimate pose and optimize it in combination with dense features.
It achieves low-cost, high-precision, and large-scale 3D scene reconstruction, ensuring the accuracy and reliability of the reconstructed 3D model in conforming to the geometric structure and spatial layout of the real scene, improving the efficiency and accuracy of 3D scene reconstruction, and solving the problems of image matching and pose effect without keypoint detection in weak texture areas, ensuring the accuracy and integrity of reconstruction.
Smart Images

Figure CN120997386A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of three-dimensional scene reconstruction, in particular to a large-scale three-dimensional scene reconstruction method, system, device and medium. BACKGROUND
[0002] As a core means of digital perception of the real world, three-dimensional reconstruction technology has been applied from professional surveying and mapping to consumer-level scenarios such as VR and AR. In the field of virtual reality (VR), according to the platform data of a certain head-mounted VR product, 72% of the three-dimensional scenes in the existing VR content rely on photogrammetry reconstruction, among which consumer-level cameras contribute 58% of the original data source.
[0003] Structure of motion (SfM) is a necessary process of three-dimensional reconstruction. Existing SfM methods are mainly divided into global methods, incremental methods and hybrid methods. The global method optimizes all two-view geometric relationships in the view graph at one time to restore the geometric information of all cameras, which has the characteristics of high efficiency and strong scalability, and is suitable for fast reconstruction of large-scale data sets, but the precision and robustness are relatively low. The incremental method starts from two views and gradually expands the reconstruction, which realizes high precision and robustness by alternately performing camera pose estimation, triangulation and bundle adjustment, but has high computational cost and cumulative error, and is suitable for high-precision scene reconstruction of small-scale data sets. The hybrid method combines the efficiency of the global method and the robustness of the incremental method, and through graph partitioning or incremental estimation, it takes into account the processing capacity of large-scale data sets and high-precision requirements, but its implementation is relatively complex, and it depends on the accuracy of the camera internal parameters.
[0004] However, most existing SfM methods rely on professional equipment for data acquisition, such as backpack mobile measurement systems or mobile measurement vehicles, which are usually equipped with high-precision sensors and positioning systems (such as GNSS, IMU, etc.), and can provide accurate initial position information and image data, thereby obtaining high-quality three-dimensional reconstruction results. However, these professional equipment is expensive and difficult to meet the needs of existing applications. In addition, existing methods are inefficient when processing large-scale data, and the reconstruction time is long, and the reconstruction fails in weak texture areas. In summary, the existing three-dimensional reconstruction technology has obvious shortcomings in professional equipment dependence, reconstruction accuracy and efficiency, which limits its application in consumer-level fast reconstruction tasks. SUMMARY
[0005] The purpose of the present application is to provide a large-scale three-dimensional scene reconstruction method, system, device and medium, which can perform three-dimensional scene reconstruction based on consumer-level equipment and improve the accuracy and efficiency of three-dimensional scene reconstruction.
[0006] To achieve the above object, the application provides a large-scale three-dimensional scene reconstruction method, comprising:
[0007] The positioning device and the image acquisition device are time-synchronized, geographical space coordinates and video frame images are acquired, the video frame images are given geographical space coordinates, key image frames are extracted based on spatial distances between video frames, video frame dense features and video frame global features, and a key image frame set is obtained;
[0008] Based on the key image frame set, three-dimensional point cloud reconstruction and image pose optimization are performed by using bundle adjustment to obtain a motion structure recovery result.
[0009] Optionally, after the motion structure recovery result is obtained by performing three-dimensional point cloud reconstruction and image pose optimization by using bundle adjustment based on the key image frame set, the method further comprises:
[0010] The information entropy of the key image frame is calculated, and a weak texture area of the video frame image is identified;
[0011] According to the weak texture area, an image matching network without key point detection is used to estimate the image pose of the weak texture area;
[0012] According to the image pose and the video frame dense feature, the motion structure recovery result is optimized by using bundle adjustment combined with the video frame dense feature according to a second target function, and an optimized motion structure recovery result is obtained.
[0013] Optionally, the time synchronization of the positioning device and the image acquisition device, the acquisition of geographical space coordinates and video frame images, the giving of geographical space coordinates to the video frame images, the extraction of key image frames based on spatial distances between video frames, video frame dense features and video frame global features, and the obtaining of a key image frame set comprise:
[0014] The positioning device and the image acquisition device are time-synchronized through a network time protocol;
[0015] The GPS data sequence of the positioning device and the video frame sequence of the image acquisition device are acquired; wherein the GPS data sequence contains time stamps and geographical space coordinates, and the video frame sequence contains time stamps;
[0016] Based on the synchronized time stamps, the geographical coordinates of each video frame are calculated by linear interpolation, and are converted into three-dimensional coordinates in the northeast celestial coordinate system;
[0017] The spatial distances are calculated based on the three-dimensional coordinates of adjacent video frames, video frame extraction is performed according to a preset spatial distance threshold, and a video frame set is generated;
[0018] extracting dense features of the video frame set and reducing dimensions by principal component analysis to obtain global features of the video frames;
[0019] screening the video frame set based on cosine similarity and spatial distance of the global features to obtain a key image frame set.
[0020] Optionally, based on the key image frame set, bundle adjustment is used for three-dimensional point cloud reconstruction and image pose optimization to obtain a motion structure recovery result, including:
[0021] using a feature point detection algorithm to extract feature points and corresponding 128-dimensional feature vectors from the key image frame set;
[0022] based on the geographic spatial coordinates of the key image frames, constructing a pair image;
[0023] based on the pair image, calculating feature distances between the extracted 128-dimensional feature vectors corresponding to the feature points;
[0024] calculating a nearest neighbor distance ratio according to the feature distances;
[0025] screening matching feature pixels according to the nearest neighbor distance ratio;
[0026] based on the matching feature pixels, generating an initial three-dimensional point cloud by essential matrix decomposition;
[0027] calculating initial pose parameters of the key image frames according to prior spatial constraints;
[0028] optimizing the initial three-dimensional point cloud and the initial pose parameters by bundle adjustment according to a first target function to obtain a motion structure recovery result.
[0029] Optionally, the first target function is:
[0030]
[0031] wherein, [R i |T o ] represents the pose of the key image frame i, X j represents the three-dimensional point cloud j, x ij is the two-dimensional pixel corresponding to the three-dimensional point X j , and K is the intrinsic matrix of the image acquisition device.
[0032] Optionally, the second target function is:
[0033]
[0034] wherein, [R i |T irepresents the pose of the key image frame i, X j represents a three-dimensional point cloud j, x ij is a three-dimensional point X j corresponding two-dimensional pixels, K is an intrinsic matrix of the image acquisition device, and F(·) represents taking the dense features of the key image frame i.
[0035] To achieve the above object, the present application further provides a large-scale three-dimensional scene reconstruction system, comprising:
[0036] An image data acquisition and preprocessing module is configured to time-synchronize the positioning device and the image acquisition device, obtain geographic spatial coordinates and video frame images, assign geographic spatial coordinates to the video frame images, extract key image frames based on spatial distances between video frames, dense features of the video frames, and global features of the video frames, and obtain a key image frame set.
[0037] A motion structure recovery module is configured to perform three-dimensional point cloud reconstruction and image pose optimization based on the key image frame set by using bundle adjustment, and obtain a motion structure recovery result.
[0038] Optionally, the large-scale three-dimensional scene reconstruction system further comprises a motion structure recovery optimization module configured to:
[0039] calculate the information entropy of the key image frames, and identify weak texture regions of the video frames;
[0040] estimate image poses of the weak texture regions by using an image matching network without key point detection according to the weak texture regions;
[0041] optimize the motion structure recovery result by using bundle adjustment combined with the dense features of the video frames according to the second target function based on the image poses and the dense features of the video frames, and obtain an optimized motion structure recovery result.
[0042] To achieve the above object, the present application further provides a large-scale three-dimensional scene reconstruction device, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement the large-scale three-dimensional scene reconstruction method according to any one of the above.
[0043] To achieve the above object, the present application further provides a computer readable storage medium storing a computer program, wherein the computer program controls a device where the computer readable storage medium is located to execute the large-scale three-dimensional scene reconstruction method according to any one of the above when the computer program runs.
[0044] The present application provides a large-scale three-dimensional scene reconstruction method, system, device and medium. Firstly, the positioning device and the image acquisition device are time-synchronized, and the video frame image is assigned with a geographic space coordinate. The time information of the geographic coordinate and the video frame can be accurately associated, the matching error of the coordinate and the image caused by time misalignment is avoided, an accurate space-time basis is provided for subsequent three-dimensional reconstruction, and the reconstruction result is ensured to meet the space-time logic of the real scene. Further, the key image frames are extracted based on the spatial distance between the video frames, the dense features and the global features. The data amount to be processed can be greatly reduced, the calculation complexity is reduced, the key information and features in the scene are retained, the interference of redundant frames is avoided, and the subsequent reconstruction process is ensured to be efficient and comprehensive in capturing the scene features. Finally, the three-dimensional point cloud reconstruction and image pose optimization are performed on the key image frame set by using the bundle adjustment, the camera pose parameters and the three-dimensional point cloud coordinates can be optimized at the same time, the error of the three-dimensional point projection to the image plane and the actual observed pixel point is minimized, the accuracy and reliability of the three-dimensional scene reconstruction are significantly improved, and the reconstructed three-dimensional model is more consistent with the geometric structure and spatial layout of the real scene. In addition, the image matching network without key point monitoring is applied to estimate the image pose of the weak texture region by identifying the weak texture region, and the image pose is further optimized based on the global motion structure recovery. The feature matching and pose estimation problem of the weak texture region can be effectively solved, the integrity of the three-dimensional reconstruction is ensured, the missing or error of the weak texture part in the scene is avoided, and the accuracy of the three-dimensional scene reconstruction is improved. BRIEF DESCRIPTION OF DRAWINGS
[0045] In order to more clearly illustrate the technical solutions of the present application, the drawings used in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0046] Figure 1 is a flow diagram of a large-scale three-dimensional scene reconstruction method provided by an embodiment of the present application;
[0047] Figure 2 is another flow diagram of a large-scale three-dimensional scene reconstruction method provided by an embodiment of the present application;
[0048] Figure 3 is a structural block diagram of a large-scale three-dimensional scene reconstruction system provided by an embodiment of the present application;
[0049] Figure 4 is a structural block diagram of a large-scale three-dimensional scene reconstruction device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0050] With reference to the drawings and embodiments of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of the present application.
[0051] With reference to the drawings and embodiments of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of the present application. Figure 1 Figure 1 is a flowchart of a large-scale three-dimensional scene reconstruction method provided by an embodiment of the present application, and the large-scale three-dimensional scene reconstruction method comprises steps S1 to S2:
[0052] S1, time synchronization is performed on a positioning device and an image acquisition device, geographical space coordinates and video frame images are acquired, the video frame images are given geographical space coordinates, key image frames are extracted based on spatial distances between video frames, dense features of video frames and global features of video frames, and a key image frame set is obtained;
[0053] It should be noted that, unlike the existing SfM method which mostly relies on professional equipment for data acquisition, such as a backpack type mobile measurement system or a mobile measurement vehicle, these devices are usually equipped with high-precision sensors and positioning systems (such as GNSS, IMU, etc.), and can provide accurate initial position information and image data. The positioning device in the embodiment of the present application can be a consumer-level device, such as a mobile phone, a tablet computer or the like with GPS positioning function. Similarly, the image acquisition device can be a consumer-level motion camera, a mobile phone, a single-lens reflex camera or the like.
[0054] The positioning device in the embodiment of the present application is mainly used to provide geographical coordinates, and the image acquisition device is only used to shoot images, and then the position / pose information of the positioning device is indirectly given to the frame images shot by the image acquisition device by using a time stamp.
[0055] In an optional embodiment, the step S1 comprises:
[0056] The positioning device and the image acquisition device are time synchronized through a network time protocol;
[0057] GPS data sequences of the positioning device and video frame sequences of the image acquisition device are acquired; wherein the GPS data sequences contain time stamps and geographical space coordinates, and the video frame sequences contain time stamps;
[0058] Based on the synchronized time stamps, the geographical coordinates of each video frame are calculated through linear interpolation, and are converted into three-dimensional coordinates in a northeast celestial coordinate system;
[0059] The spatial distance is calculated based on three-dimensional coordinates of adjacent video frames, video frame extraction is performed according to a preset spatial distance threshold, and a video frame set is generated;
[0060] Dense features of the video frame set are extracted, and global features of the video frames are obtained by principal component analysis dimension reduction;
[0061] The video frame set is screened based on cosine similarity and spatial distance of the global features, and a key image frame set is obtained.
[0062] S2, based on the key image frame set, three-dimensional point cloud reconstruction and image pose optimization are performed by bundle adjustment to obtain a motion structure recovery result.
[0063] In an optional embodiment, the step S2 comprises:
[0064] Feature points and corresponding 128-dimensional feature vectors are extracted from the key image frame set using a feature point detection algorithm;
[0065] Based on the geographic spatial coordinates of the key image frames, a pair of images is constructed;
[0066] Based on the pair of images, the feature distance between the extracted 128-dimensional feature vectors corresponding to the feature points is calculated;
[0067] The nearest neighbor distance ratio is calculated according to the feature distance;
[0068] The matching feature pixel points are screened according to the nearest neighbor distance ratio;
[0069] Based on the matching feature pixel points, an initial three-dimensional point cloud is generated by essential matrix decomposition;
[0070] The initial pose parameters of the key image frames are calculated according to the prior spatial constraints;
[0071] According to a first target function, the initial three-dimensional point cloud and the initial pose parameters are optimized by bundle adjustment to obtain a motion structure recovery result.
[0072] Specifically, the first target function is:
[0073]
[0074] Where [R i |T i ] represents the pose of the key image frame i, X j represents the three-dimensional point cloud j, x ij is the two-dimensional pixel corresponding to the three-dimensional point X j , and K is the intrinsic matrix of the image acquisition device.
[0075] Exemplarily, a smartphone with GPS positioning function can be selected as the positioning device, and a GoPro consumer-level sports camera can be selected as the image acquisition device. Through the NTP protocol, the clocks of the smartphone and the sports camera are synchronized to ensure consistency in time. The staff holds the sports camera and slowly moves along the target shooting scene while shooting, and the smartphone records the corresponding GPS data sequence, including the timestamp, longitude, latitude, and altitude. The sports camera acquires a video frame sequence, each of which has a timestamp. Then, according to the synchronized timestamps, the two GPS points closest to the video frame timestamp are located, the geographic coordinates of the video frame are calculated using linear interpolation, and the three-dimensional coordinates in the East-North-Up (ENU) coordinate system are converted. Then, the spatial distance of adjacent video frames in the ENU coordinate system is calculated, and the video frames are extracted according to a certain spatial distance threshold. The dense features of the extracted video frame set are extracted, and the global features of the video frame image are obtained by principal component analysis dimension reduction. Based on the cosine similarity and spatial distance of the global features, a set of key image frames is selected. For example, if the cosine similarity threshold is set to 0.6 and the spatial distance threshold is set to 5 meters, when the cosine similarity of adjacent video frames is less than 0.6 and the spatial distance is greater than 5 meters, the video frame is determined as a key image frame.
[0076] Then, the SIFT algorithm is used to extract feature points and corresponding 128-dimensional feature vectors from the key image frames. Based on the geographic coordinates of the key image frames, image pairs with possible overlapping fields of view are constructed, the feature distances between the SIFT descriptors of the feature pixel points of the image pairs are calculated, the successfully matched feature pixel point pairs are selected based on the nearest distance ratio, the essential matrix is calculated based on the matched feature pixel point pairs, the relative transformation between the key image frames is recovered by SVD decomposition of the essential matrix, the corresponding three-dimensional points are calculated using triangulation, and a three-dimensional point set is formed. The GPS coordinates obtained by the smartphone are combined as prior spatial constraints to verify and constrain the direction, scale, and position of the camera pose recovered by the two-view geometry. By eliminating the gross errors that do not match the prior space, more reliable initial poses and three-dimensional point clouds are provided for subsequent global SfM reconstruction. The bundle adjustment is used to optimize all parameters, including the camera pose (rotation matrix R and translation vector T) and the coordinates of the three-dimensional points. The objective function of the optimization is to minimize the re-projection error, i.e., the sum of the squared errors between the projected points and the actual observed points. Through this step, the three-dimensional point cloud reconstruction and image pose optimization are realized, and the motion structure recovery result is obtained, presenting a three-dimensional point cloud model of the target scene, including the spatial positions and shapes of the main elements such as buildings, roads, and greenery.
[0077] In summary, the large-scale three-dimensional scene reconstruction method provided by the embodiment of the application firstly synchronizes the positioning device and the image acquisition device in time, and assigns geographical space coordinates to video frame images, so that the time information of the geographical coordinates and the video frames can be accurately associated, the matching error of the coordinates and the images caused by time misalignment is avoided, an accurate space-time basis is provided for subsequent three-dimensional reconstruction, and the reconstruction result is ensured to conform to the space-time logic of the real scene. Further, the key image frames are extracted based on the spatial distance between the video frames, the dense features and the global features, so that the amount of data to be processed can be greatly reduced, the calculation complexity is reduced, the key information and features in the scene can be retained, the interference of redundant frames is avoided, and it is ensured that the subsequent reconstruction process efficiently and comprehensively captures the scene features. Finally, the bundle adjustment is used to perform three-dimensional point cloud reconstruction and image pose optimization on the key image frame set, so that the camera pose parameters and the three-dimensional point cloud coordinates can be optimized at the same time, the error of the three-dimensional point projection to the image plane and the actual observed pixel point is minimized, the accuracy and reliability of the three-dimensional scene reconstruction are significantly improved, and the reconstructed three-dimensional model is more consistent with the geometric structure and spatial layout of the real scene.
[0078] In an optional embodiment, after the step S2, the method further comprises:
[0079] calculating the information entropy of the key image frame, and identifying the weak texture area of the video frame image;
[0080] According to the weak texture area, the image matching network without key point detection is used to estimate the image pose of the weak texture area.
[0081] According to the image pose and the dense features of the video frame, the motion structure recovery result is optimized by using the bundle adjustment combined with the dense features of the video frame according to a second target function, to obtain an optimized motion structure recovery result.
[0082] Specifically, the second target function is:
[0083]
[0084] wherein, [R i |T i ] represents the pose of the key image frame i, X j represents the three-dimensional point cloud j, x ij is the two-dimensional pixel corresponding to the three-dimensional point X j , and K is the intrinsic matrix of the image acquisition device, and F(·) represents the dense features of the key image frame i.
[0085] It is worth mentioning that the embodiment of the application firstly constructs relative constraints based on key images to recover global motion structure, then identifies weak texture regions (regions with unobvious texture features, single visual information or lack of significant structural changes in images), applies an image matching network without key point monitoring to estimate the image pose of the weak texture region, and further optimizes the image pose by applying dense features on the basis of global motion structure recovery. The two-stage aerial triangulation framework is innovatively designed, which can effectively solve the feature matching and pose estimation problem of such regions, ensure the integrity of three-dimensional reconstruction, avoid the missing or error of weak texture part in the scene, and realize the collaborative improvement from the whole to the local on the basis of the overall framework of global motion structure recovery. The global structure constraint is used to ensure the correctness of the reconstruction direction, and the local dense feature optimization is used to fill in the details and correct the deviation, so as to finally improve the quality and accuracy of the whole three-dimensional scene reconstruction. The embodiment of the application can realize large-scale global motion structure recovery by using low-cost sensors such as mobile phones and consumer-grade motion cameras, and can realize more low-cost and large-scale three-dimensional scene reconstruction by combining the network time synchronization to assign geographical space coordinates to images and the two-stage efficient motion structure recovery framework.
[0086] In order to make the person skilled in the art more clear about the implementation process of the large-scale three-dimensional scene reconstruction method provided by the embodiment of the application, the following will take a mobile phone as a positioning device and a motion camera as an image acquisition device as an example to make a more detailed description.
[0087] Referring to Figure 2 , Figure 2 is another flowchart of the large-scale three-dimensional scene reconstruction method provided by the embodiment of the application. As shown in Figure 2 , specifically, the large-scale three-dimensional scene reconstruction method comprises the following steps 1 to 15:
[0088] Step 1: trajectory and image data acquisition: time synchronization of the mobile phone and the motion camera is performed through a network time protocol (such as NTP protocol), then a GPS data sequence {G k} of the mobile phone is recorded, wherein G k =(t k ,lon k ,lat k ,alt k ), and a video frame sequence {F m} of the motion camera is obtained, wherein each frame contains a time stamp t m .
[0089] It should be noted that t k represents the time corresponding to the kth data in the GPS data sequence recorded by the mobile phone. lon kis the abbreviation of "longitude", which represents the longitude value of the location where the mobile phone is located at the kth recording moment, and is used to indicate the coordinates of the location in the east-west direction, with the prime meridian as the reference, east longitude (value range 0°-180°) to the east, and west longitude (value range 0°-180°) to the west. lat k is the abbreviation of "latitude", which represents the latitude value of the location where the mobile phone is located at the kth recording moment, and is used to indicate the coordinates of the location in the north-south direction, with the equator as the reference, north latitude (value range 0°-90°) to the north, and south latitude (value range 0°-90°) to the south. alt k is the abbreviation of "altitude", which represents the altitude value of the location where the mobile phone is located at the kth recording moment, that is, the vertical height of the location from the average sea level, and the unit is usually meters.
[0090] Step 2: Video frame geographic coordinate calculation: for any video frame F m , locate the nearest GPS points G m (t a ) and G a (t b ) before and after the timestamp t b , calculate the geographic spatial coordinates by linear interpolation:
[0091]
[0092] Convert to northeast celestial coordinate system (x m , y m , z m ), and the coordinate conversion formula is:
[0093]
[0094] In the formula, T w is the conversion matrix from the WGS84 coordinate system to the ENU coordinate system, and (lon0, lat0, alt0) is the origin of the local coordinate system.
[0095] Step 3: Video frame spatial sampling: calculate the spatial distance between adjacent video frames, and perform video frame extraction according to a certain spatial distance to generate a video frame set {F s}, which is used for subsequent three-dimensional reconstruction.
[0096] Step 4: Motion camera intrinsic parameter calibration: use the chessboard calibration method to calibrate the intrinsic parameters of the motion camera, and obtain the intrinsic matrix K of the video frames in the video frame set {F s}:
[0097]
[0098] where K is a 3x3 matrix, f x , f y represent the focal length of video frame in X, Y direction respectively, c x , c y represent the pixel coordinate of video frame in X, Y direction respectively.
[0099] Step 5: Video frame dense feature and global feature extraction: Extract dense feature D s from video frame set {F s} and get global feature g s by principal component analysis (PCA) dimension reduction.
[0100] Step 6, key image frame extraction: Based on video frame global feature g s , calculate the similarity between global features by cosine distance:
[0101]
[0102] where s ij represents the similarity between global features, the smaller s ij is, the lower the similarity between images is.
[0103] Keep the adjacent images with s ij less than the threshold value as key image frames, and get the key image frame set.
[0104] So far, the above steps 1-6 complete the collection and preprocessing of image data.
[0105] Step 7, feature point extraction: Use feature point detection algorithm SIFT to extract feature points (u i , v i ) and corresponding 128-dimensional feature vectors from key images in key image frame set.
[0106] Step 8, image pairing: Based on the geographic coordinates of key image frames, construct image pairs {(F p , F q )} that may have overlapping fields of view.
[0107] Step 9, feature point matching: For the paired images in step 8, calculate the feature distance between the 128-dimensional feature vectors corresponding to the feature points (u i , v i ) extracted in step 7.
[0108] Based on the nearest distance ratio NNDR, filter the successfully matched feature pixel point pairs:
[0109]
[0110] In the formula: dist1 is the minimum feature distance, dist2 is the second smallest feature distance, and if NNDR < thr, it is determined that the feature pixel point pair is a successful matching. Based on the random sample consensus algorithm, the basis matrix F is estimated, and the inner point set is reserved.
[0111] Step 10, double-view geometry recovery: based on the feature pixel point pair, the essential matrix F is calculated:
[0112] F = K -T FK -1 ;
[0113] The relative conversion (R, T) is recovered by SVD decomposition of the essential matrix F, and the corresponding three-dimensional point X is calculated using triangulation:
[0114]
[0115] In the formula, P1 = [I|0], P2 = [R|T], the three-dimensional point set {X n} is formed by reserving the point X with positive depth and minimum re-projection error.
[0116] Based on the conversion calculated based on the geographical coordinates between the key frames and the relative conversion obtained by double-view geometry recovery, the random sample consensus algorithm is used to verify the relative translation direction and scale, and the gross error is removed.
[0117] It should be noted that the verification and constraint of the camera pose recovered by double-view geometry in terms of direction, scale and position are performed by removing the gross error matching inconsistent with the prior space, so as to provide more reliable initial pose and three-dimensional point cloud for subsequent global SfM reconstruction, and finally realize high-precision three-dimensional scene reconstruction under low-cost equipment. This design ingeniously combines the positioning ability of consumer-level equipment and geometric vision algorithm, and makes up for the dependence of traditional SfM on professional equipment.
[0118] Step 11, image pose initialization: after selecting the reference coordinate system, the absolute pose (R k ,T k ) is calculated through the 3D-2D point correspondence:
[0119] λx = K[R k |T k ]X;
[0120] In the formula, λ is the depth factor, and K is the image intrinsic matrix.
[0121] It should be noted that the core role of step 11 is to convert the relative pose recovered by the dual-view geometry into a globally unified absolute attitude through the 3D-2D point correspondence, so as to initialize the pose parameters for the subsequent bundle adjustment, so that the bundle adjustment can optimize the poses and three-dimensional point clouds of all frames simultaneously.
[0122] Step 12, global attitude optimization: all parameters are optimized by bundle adjustment to obtain the motion structure recovery result, including camera pose parameters and three-dimensional point cloud data. The objective function is as follows:
[0123]
[0124] In the formula, [R i |T i ] is the pose of image i, x ij is the corresponding two-dimensional pixel of three-dimensional point X j .
[0125] Up to now, the above steps 7-12 obtain a one-stage motion structure recovery result.
[0126] Step 13, weak texture area detection: calculate the information entropy of the key frame image to determine the weak texture area. Information entropy is an index for measuring information uncertainty. When applied to images, it reflects the complexity of pixel gray scale distribution. The specific calculation method is as follows:
[0127]
[0128] In the formula, p i represents the frequency of gray level appearing in the image, which is obtained by histogram statistics.
[0129] Step 14, weak texture area image extraction: extract all images near the weak texture area image from all images obtained in step 3 to enhance image density. Use the detector-free dense / semi-dense matcher LoFTR to establish dense correspondence between image pairs directly through attention mechanism. After establishing the correspondence, perform steps 9 to 10 to calculate the fine pose of the weak texture area image and obtain more fine image pose.
[0130] Step 15, global optimization combined with dense features: based on the image pose recovered in step 14 and the motion structure recovery result output in step 12, perform global bundle adjustment to further optimize the camera's internal and external parameters and the position of three-dimensional points, and improve the accuracy and quality of three-dimensional reconstruction. The global bundle adjustment combines the dense feature information of the image and the re-projection error to optimize the parameters, and obtains the optimized motion structure recovery result, which can ensure the accuracy and consistency of the optimization result. The objective function of this time is as follows:
[0131]
[0132] In the formula, F(·) represents taking the dense features of the key image frame i.
[0133] Up to now, the one-stage motion structure recovery result is optimized by the steps 13-15, and the optimized two-stage motion structure recovery result is obtained.
[0134] In summary, the large-scale three-dimensional scene reconstruction method provided by the embodiment of the application firstly assigns geographical space coordinates to images formed by frame extraction of a video based on network time synchronization mobile phones and motion cameras, designs a key frame global SfM method based on prior spatial constraints, realizes fast one-stage reconstruction, and on this basis, identifies weak texture regions based on image information entropy, estimates weak texture region image poses by using a key point detection-free image matching network, solves the problem of poor reconstruction accuracy of weak texture regions, and on the basis of global motion structure recovery, further optimizes the SfM result by using dense features. In addition, geographical space coordinates are assigned to images based on network time synchronization, and a two-stage efficient SfM method is combined to support lower-cost and larger-range three-dimensional scene reconstruction, and the accuracy and efficiency of three-dimensional scene reconstruction are improved.
[0135] Referring to Figure 3 , Figure 3 is a structural block diagram of a large-scale three-dimensional scene reconstruction system 200 provided by the embodiment of the application, and the large-scale three-dimensional scene reconstruction system 200 comprises:
[0136] An image data acquisition and preprocessing module 21 is configured to synchronize a positioning device and an image acquisition device in time, acquire geographical space coordinates and video frame images, assign geographical space coordinates to the video frame images, extract key image frames based on spatial distances between video frames, dense features of video frames, and global features of video frames, and obtain a key image frame set.
[0137] A motion structure recovery module 22 is configured to perform three-dimensional point cloud reconstruction and image pose optimization by using bundle adjustment based on the key image frame set, and obtain a motion structure recovery result.
[0138] In an optional embodiment, the large-scale three-dimensional scene reconstruction system 200 further comprises a motion structure recovery optimization module 23 configured to:
[0139] Calculate information entropy of the key image frames, and identify weak texture regions of the video frame images;
[0140] Estimate weak texture region image poses by using a key point detection-free image matching network according to the weak texture regions;
[0141] According to the image pose and the video frame dense feature, an optimized motion structure recovery result is obtained by optimizing the motion structure recovery result according to a second target function by using bundle adjustment combined with the video frame dense feature.
[0142] In an optional embodiment, the image data acquisition and preprocessing module 21 is specifically configured to:
[0143] Time synchronization is performed on the positioning device and the image acquisition device through a network time protocol;
[0144] A GPS data sequence of the positioning device and a video frame sequence of the image acquisition device are acquired, wherein the GPS data sequence comprises time stamps and geographic spatial coordinates, and the video frame sequence comprises time stamps;
[0145] Based on the synchronized time stamps, the geographic coordinates of each video frame are calculated through linear interpolation, and are converted into three-dimensional coordinates in the northeast celestial coordinate system;
[0146] The spatial distances are calculated based on the three-dimensional coordinates of adjacent video frames, video frames are extracted according to a preset spatial distance threshold, and a video frame set is generated;
[0147] Dense features of the video frame set are extracted, and global features of the video frames are obtained through principal component analysis dimension reduction;
[0148] The video frame set is screened based on the cosine similarity of the global features and the spatial distances, and a key image frame set is obtained.
[0149] In an optional embodiment, the motion structure recovery module 22 is specifically configured to:
[0150] Feature points and corresponding 128-dimensional feature vectors are extracted from the key image frame set using a feature point detection algorithm;
[0151] Based on the geographic spatial coordinates of the key image frames, paired images are constructed;
[0152] Based on the paired images, feature distances between the extracted 128-dimensional feature vectors of the feature points are calculated;
[0153] A nearest neighbor distance ratio is calculated according to the feature distances;
[0154] Matching feature pixel points are screened according to the nearest neighbor distance ratio;
[0155] Based on the matching feature pixel points, an initial three-dimensional point cloud is generated through essential matrix decomposition;
[0156] Initial pose parameters of the key image frames are calculated according to prior spatial constraints;
[0157] According to the first target function, the initial three-dimensional point cloud and the initial pose parameter are optimized by using bundle adjustment to obtain a motion structure recovery result.
[0158] It should be noted that the large-scale three-dimensional scene reconstruction system provided by the embodiments of the present application is used to execute all process steps of the large-scale three-dimensional scene reconstruction method provided by the above embodiments, and the working principles and beneficial effects of the two are one-to-one correspondence, thus not being repeated.
[0159] Referring to Figure 4 , Figure 4 is a structural block diagram of a large-scale three-dimensional scene reconstruction device 300 provided by the embodiments of the present application, the large-scale three-dimensional scene reconstruction device 300 includes a processor 31, a memory 32, and a computer program stored in the memory 32 and executable on the processor 31. The processor 31 implements the steps in each of the above large-scale three-dimensional scene reconstruction method embodiments when executing the computer program, such as steps S1-S3.
[0160] Exemplarily, the computer program can be divided into one or more modules / units, which are stored in the memory 32 and executed by the processor 31 to complete the present application. The one or more modules / units can be a series of computer program instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program in the large-scale three-dimensional scene reconstruction device 300.
[0161] The large-scale three-dimensional scene reconstruction device 300 can include, but is not limited to, the processor 31, the memory 32. Those skilled in the art can understand that the schematic diagram is only an example of the large-scale three-dimensional scene reconstruction device 300, and does not constitute a limitation on the large-scale three-dimensional scene reconstruction device 300, and can include more or fewer components than the diagram, or combine certain components, or different components, for example, the large-scale three-dimensional scene reconstruction device 300 can also include an input / output device, a network access device, a bus, etc.
[0162] The processor 31 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic components, discrete hardware components, or the like. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The processor 31 is a control center of the large-scale three-dimensional scene reconstruction device 300, and is connected to various parts of the large-scale three-dimensional scene reconstruction device 300 through various interfaces and lines.
[0163] The memory 32 can be used to store computer programs and / or modules. The processor 31 realizes various functions of the large-scale three-dimensional scene reconstruction device 300 by running or executing computer programs and / or modules stored in the memory 32, and calling data stored in the memory 32. The memory 32 can mainly include a program storage area and a data storage area. The program storage area can store an operating system, at least one application program required for a function (such as a sound playing function, an image playing function, etc.), and the like. The data storage area can store data created according to use of the mobile phone (such as audio data, a phone book, etc.), and the like. In addition, the memory 32 can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state memory device.
[0164] The modules / units integrated by the large-scale three-dimensional scene reconstruction device 300 can be stored in a computer readable storage medium if they are realized in the form of software function units and sold or used as independent products. Based on this understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by a computer program instructing related hardware. The computer program can be stored in a computer readable storage medium. When the computer program is executed by the processor 31, the steps of the above-mentioned various method embodiments can be realized. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium, etc.
[0165] The above is the preferred embodiment of the present application. It should be pointed out that, for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, which are also considered within the scope of protection of the present application.
Claims
1. A method for large-scale 3D scene reconstruction, characterized in that, include: The positioning device and the image acquisition device are synchronized in time to obtain geospatial coordinates and video frame images. Geospatial coordinates are assigned to the video frame images. Key image frames are extracted based on the spatial distance between video frames, the density features of video frames, and the global features of video frames to obtain a set of key image frames. Based on the set of key image frames, bundle adjustment is used to reconstruct 3D point clouds and optimize image pose, resulting in motion structure recovery.
2. The large-scale three-dimensional scene reconstruction method as described in claim 1, characterized in that, After performing 3D point cloud reconstruction and image pose optimization using bundle adjustment based on the key image frame set to obtain the motion structure recovery result, the method further includes: Calculate the information entropy of key image frames to identify weak texture regions in video frame images; Based on the weak texture region, an image matching network without keypoint detection is used to estimate the pose of the weak texture region image. Based on the image pose and the dense features of the video frames, and according to the second objective function, the motion structure restoration result is optimized by using bundle adjustment that combines the dense features of the video frames, resulting in an optimized motion structure restoration result.
3. The large-scale three-dimensional scene reconstruction method as described in claim 1, characterized in that, The process involves synchronizing the positioning device and the image acquisition device in time to obtain geospatial coordinates and video frame images. Geospatial coordinates are assigned to the video frame images. Key image frames are extracted based on the spatial distance between video frames, video frame density features, and global features of the video frames, resulting in a set of key image frames, including: The positioning device and the image acquisition device are synchronized in time using the Network Time Protocol (NTP). The GPS data sequence of the positioning device and the video frame sequence of the image acquisition device are collected; wherein, the GPS data sequence includes timestamps and geospatial coordinates, and the video frame sequence includes timestamps; Based on the synchronized timestamp, the geographic coordinates of each video frame are calculated by linear interpolation and converted into three-dimensional coordinates in the northeast-northeast coordinate system. The spatial distance is calculated based on the three-dimensional coordinates of adjacent video frames, and video frames are extracted according to a preset spatial distance threshold to generate a set of video frames. Dense features of the video frame set are extracted, and global features of the video frames are obtained by dimensionality reduction through principal component analysis; The set of video frames is filtered based on the cosine similarity and spatial distance of the global features to obtain a set of key image frames.
4. The large-scale three-dimensional scene reconstruction method as described in claim 1, characterized in that, Based on the set of key image frames, 3D point cloud reconstruction and image pose optimization are performed using bundle adjustment to obtain motion structure recovery results, including: Feature points and corresponding 128-dimensional feature vectors are extracted from the key image frame set using a feature point detection algorithm. Based on the geospatial coordinates of key image frames, paired images are constructed; Based on the paired images, the feature distance between the 128-dimensional feature vectors corresponding to the extracted feature points is calculated; Calculate the nearest neighbor distance ratio based on the aforementioned characteristic distance; Filter matching feature pixels based on the nearest neighbor distance ratio; Based on the matched feature pixels, an initial 3D point cloud is generated through essential matrix decomposition. Calculate the initial pose parameters of the key image frame based on prior spatial constraints; Based on the first objective function, the initial 3D point cloud and the initial pose parameters are optimized using bundle adjustment to obtain the motion structure recovery result.
5. The large-scale three-dimensional scene reconstruction method as described in claim 4, characterized in that, The first objective function is: Among them, [R i |T i ] represents the pose of keyframe i, X j Represents a 3D point cloud j, x ij For a 3D point X j The corresponding two-dimensional pixels, K is the intrinsic parameter matrix of the image acquisition device.
6. The large-scale three-dimensional scene reconstruction method as described in claim 2, characterized in that, The second objective function is: Among them, [R i |T i ] represents the pose of keyframe i, X j Represents a 3D point cloud j, x ij For a 3D point X j The corresponding two-dimensional pixels, K is the intrinsic parameter matrix of the image acquisition device, and F(·) represents the dense features of the key image frame i.
7. A large-scale three-dimensional scene reconstruction system, characterized in that, include: The image data acquisition and preprocessing module is used to synchronize the positioning device and the image acquisition device in time, acquire geospatial coordinates and video frame images, assign geospatial coordinates to the video frame images, and extract key image frames based on the spatial distance between video frames, the density features of video frames and the global features of video frames to obtain a set of key image frames. The motion structure recovery module is used to perform three-dimensional point cloud reconstruction and image pose optimization based on the key image frame set using bundle adjustment to obtain motion structure recovery results.
8. The large-scale three-dimensional scene reconstruction system as described in claim 7, characterized in that, The large-scale 3D scene reconstruction system also includes a motion structure recovery and optimization module, used for: Calculate the information entropy of key image frames to identify weak texture regions in video frame images; Based on the weak texture region, an image matching network without keypoint detection is used to estimate the pose of the weak texture region image. Based on the image pose and the dense features of the video frames, and according to the second objective function, the motion structure restoration result is optimized by using bundle adjustment that combines the dense features of the video frames, resulting in an optimized motion structure restoration result.
9. A large-scale three-dimensional scene reconstruction device, characterized in that, include: A processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements the large-scale three-dimensional scene reconstruction method as described in any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to perform the large-scale three-dimensional scene reconstruction method as described in any one of claims 1 to 6.