3D scene reconstruction method and system based on adaptive key frame

By dynamically adjusting keyframe selection in the SLAM system, combining visual inertial joint optimization and 3D Gaussian distributed rendering, the problem of excessive keyframes in complex scenes and insufficient in simple scenes is solved, and efficient and adaptive 3D scene reconstruction is achieved.

CN120147541AActive Publication Date: 2025-06-13江淮前沿技术协同创新中心

Patent Information

Application Number
CN202510257622.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-13
Estimated Expiration
2045-03-05

AI Technical Summary

Technical Problem

The existing SLAM system fails to fully consider the scene complexity and rendering requirements when selecting keyframes, resulting in too many keyframes in complex scenarios and insufficient keyframes in simple scenarios, which affects system performance and reconstruction quality.

Method used

By obtaining the continuous multi-frame image and inertial measurement unit data of the current scene, the target feature points are extracted and tracked, the keyframe selection is dynamically adjusted based on the scene complexity index and the system resource utilization index, and visual inertial joint optimization is performed to generate a 3D Gaussian distribution and render it.

Benefits of technology

It realizes automatic adjustment of keyframes according to scene complexity and hardware conditions, improves the adaptability and efficiency of 3D scene reconstruction, maintains efficient operation in different environments, and improves rendering quality and real-time performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147541A_ABST
    Figure CN120147541A_ABST
Patent Text Reader

Abstract

The invention provides a 3D scene reconstruction method and system based on adaptive key frames, and relates to the technical field of computer vision, and the method comprises the steps: obtaining continuous multi-frame images and inertial measurement unit data of a current scene; acquiring camera trajectory data and offset data of each inertial measurement assembly based on the inertial measurement unit data; carrying out target feature extraction and target tracking on each frame in the multiple frames of images to obtain feature point data of the to-be-tracked target; determining a key frame based on a scene complexity index and a system resource utilization rate index of the current scene; performing visual inertia joint optimization based on all feature point data, camera trajectory data and offset data, and generating point cloud data containing each feature point; and generating 3D Gaussian distribution of the current scene based on the point cloud data and the key frame, and rendering the distribution to obtain a first 3D rendered image of the current scene. According to the method, real-time 3D scene reconstruction can be carried out in a complex and changeable application scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer vision technology, and particularly to a 3D scene reconstruction method and system based on adaptive key frames. Background Art

[0002] In recent years, 3D scene reconstruction technology and Simultaneous Localization and Mapping (SLAM) technology have made remarkable progress in the fields of computer vision, augmented reality, and robot navigation. Existing SLAM systems mainly rely on feature point extraction and matching, such as Simultaneous Localization and Mapping based on Fast Feature Point Extraction and Description Algorithm (ORB-SLAM), Monocular Visual Inertial State Estimator (VINS-Mono), etc. Although they perform well in terms of positioning accuracy, it is difficult to generate high-quality dense 3D reconstruction results.

[0003] To address the limitations of sparse reconstruction, a series of methods based on the direct method, such as Monocular Simultaneous Localization and Mapping based on the Direct Method (LSD-SLAM) and Visual Odometry based on Sparse Direct Method (DSO), have been proposed. These methods directly utilize the pixel information of images for optimization and can generate semi-dense reconstruction results. However, they have low computational efficiency when dealing with large-scale scenes and are too sensitive to light changes.

[0004] Neural implicit representation methods, such as Neural Radiance Fields (NeRF), can generate high-quality novel view rendering results, but the training and rendering processes have high computational costs and it is difficult to achieve real-time performance. Currently, combining efficient neural rendering technology with real-time SLAM systems mainly faces the following challenges: First, the selection of key frames. Existing SLAM systems select key frames based on geometric information, failing to fully consider scene complexity and rendering requirements, resulting in too many key frames in complex scenes and insufficient key frames in simple scenes, affecting system performance and reconstruction quality. Second, computational resource allocation. Existing SLAM systems allocate a large amount of computational resources to feature extraction and matching at the SLAM front end, while ignoring the resource requirements of the back-end optimization and rendering processes, leading to unbalanced system performance. Third, the trade-off between rendering quality and real-time performance. High-quality 3D rendering usually requires a large amount of computational resources, which conflicts with the real-time requirements of SLAM systems and it is difficult to meet real-time requirements while ensuring rendering quality. Fourth, system robustness. In complex environments with dynamic and drastic light changes, the stability and accuracy of existing SLAM systems are often difficult to guarantee. Summary of the Invention

[0005] The main objective of the present disclosure is to provide a 3D scene reconstruction method and system based on adaptive key frames to solve the technical problem that the real-time 3D scene reconstruction effect is limited in related technologies due to the inability to adaptively adjust key frames according to scene complexity and hardware conditions.

[0006] To achieve the above object, a first aspect of the present disclosure provides a 3D scene reconstruction method based on adaptive key frames, including:

[0007] Obtain a plurality of consecutive frames of images of the current scene and inertial measurement unit data, wherein the inertial measurement unit is disposed inside the camera that captures the images and is equipped with a plurality of inertial measurement components;

[0008] Extract target features and perform target tracking on each of the plurality of frames of images to obtain feature point data of the target to be tracked;

[0009] Obtain camera trajectory data and bias data of each of the inertial measurement components based on the inertial measurement unit data;

[0010] Determine key frames based on the scene complexity index and system resource utilization index of the current scene;

[0011] Perform visual-inertial joint optimization based on all the feature point data, the camera trajectory data, and the bias data to generate point cloud data including feature points of the target to be tracked, wherein the visual-inertial joint optimization is configured to eliminate the visual reprojection error of the point cloud data by correcting the feature point data, the camera trajectory data, and the bias data; and,

[0012] Generate a 3D Gaussian distribution of the current scene based on the point cloud data and the key frames, and render the 3D Gaussian distribution to obtain a first 3D rendered image of the current scene.

[0013] Further, the method further includes:

[0014] Identify a similar scene in the historical scene that matches the current scene;

[0015] Obtain the camera pose calculation model, and perform global consistency optimization on the camera pose calculation model based on the camera trajectory data of the similar scene and the camera trajectory data of the current scene;

[0016] Obtain optimized camera trajectory data of the current scene based on the optimized camera pose calculation model; and,

[0017] Perform the visual-inertial joint optimization based on the optimized camera trajectory data, and further generate a second 3D rendered image of the current scene.

[0018] Further, the method further includes:

[0019] Generate a first sub-map based on the point cloud data of the current scene;

[0020] Identify a second scenario in the historical scenario that forms a global map closed-loop with the current scenario;

[0021] Obtain the point cloud data of the second scenario, and generate a second sub-map based on the point cloud data of the second scenario;

[0022] Merge the first sub-map and the second sub-map to generate a global map.

[0023] Further, the determination of the key frame based on the scenario complexity index and the system resource utilization index of the current scenario includes:

[0024] For each frame in the multiple frames of images as the current frame, obtain the number of feature points N f , the variance V of the feature point distribution f , the texture complexity T f ;

[0025] Based on the number of feature points N f , the variance V of the feature point distribution f , the texture complexity T f Obtain the scenario complexity index C of the current frame;

[0026] Determine the nearest key frame, and obtain the relative displacement A, the overlap degree B, and the information entropy H of the current frame relative to the nearest key frame;

[0027] Obtain the system resource utilization index R, and dynamically adjust the selection threshold T based on the scenario complexity index C and the system resource utilization index R;

[0028] Determine whether the current frame is the key frame based on the information entropy H, the selection threshold T, the relative displacement A, and the overlap degree B.

[0029] Further, the scenario complexity index C, the information entropy H, and the selection threshold T are respectively configured as:

[0030] C = w 1 ×N f +w 2 ×V f +w 3 ×T f

[0031]

[0032] T = T 0 ×(1 + α×C)×(1 - β×R)

[0033] In the formula, w 1 , w2 , w 3 is the weight coefficient, p(x i ) is the probability distribution of the feature points in the image region x i , T 0 is the basic threshold, α is the adjustment coefficient, and β is the correction coefficient; and,

[0034] The key frame is configured to simultaneously satisfy: the information entropy H exceeds the selection threshold T, and the relative displacement A exceeds a preset maximum allowable value, and the overlap degree B is lower than a preset overlap threshold.

[0035] Further, performing target feature extraction and target tracking on each of the multiple frames of images respectively to obtain feature point data of the target to be tracked, including:

[0036] For each frame of the multiple frames of images as the current frame, performing feature tracking processing on the current frame through the KLT sparse optical flow algorithm to determine the first type of feature points of the target to be tracked, and obtaining the feature point data of the first type of feature points, where the feature point data is configured to include feature point position data and feature point gray data;

[0037] Performing corner feature search on the current frame through the FAST corner detector to determine the second type of feature points of the target to be tracked, and obtaining the feature point data of the second type of feature points;

[0038] Obtaining the minimum pixel interval between adjacent feature points in the feature point set composed of the first type of feature points and the second type of feature points;

[0039] Performing 2D feature distortion removal processing on the elements of the feature point set based on the minimum pixel interval, and projecting the feature point data corresponding to the elements of the processed feature point set onto the unit sphere;

[0040] Performing outlier rejection on the projected feature point data through the random sample consensus algorithm to obtain the final feature point data of the target to be tracked.

[0041] Further, the camera trajectory data is configured to include camera position data, camera speed data, and camera attitude data, and the inertial measurement component is configured to include an accelerometer and a gyroscope; and obtaining the camera trajectory data and the bias data of each of the inertial measurement components based on the inertial measurement unit data includes:

[0042] Obtaining camera trajectory data based on the inertial measurement unit data;

[0043] Obtain the measurement model of the inertial measurement unit, and within the time interval between two consecutive image frames, perform pre-integration processing on the inertial measurement unit data based on the measurement model to obtain the relative displacement, relative velocity, and relative rotation angle of the camera within the time interval;

[0044] Determine the covariance propagation matrix of the measurement model based on the relative displacement, relative velocity, and relative rotation angle;

[0045] Optimize the parameters of the measurement model based on the covariance propagation matrix, so that the bias data of the inertial measurement components all converge to their respective set threshold ranges, and end the optimization;

[0046] Obtain the bias data of each inertial measurement component based on the measurement model after parameter optimization.

[0047] Further, the method further includes: when it is recognized that the bias data of any inertial measurement component exceeds the corresponding set threshold, calibrate the relative displacement, relative velocity, and relative rotation angle based on the measurement model, and then calibrate the camera trajectory data; and,

[0048] The visual-inertial joint optimization based on all the feature point data, the camera trajectory data, and the bias data to generate point cloud data including each feature point of the target to be tracked includes:

[0049] For each feature point of the target to be tracked as the current feature point, obtain the state data corresponding to the current feature point in each frame of the multi-frame images, where the state data is configured to include the feature point position data, the feature point gray data, the camera position data, the camera speed data, the camera attitude data, and the bias data of each inertial measurement component;

[0050] Obtain the inverse depth parameter corresponding to the current feature point in each frame of the multi-frame images, and generate the state vector X of the current feature point based on the state data and the inverse depth parameter;

[0051] Input the state vector X into the visual reprojection error model to obtain the visual reprojection error;

[0052] Input the state vector X into the pre-integration residual model to obtain the pre-integration residual;

[0053] Perform non-linear optimization on the state vector X based on the visual reprojection error and the pre-integration residual to obtain the optimized state data corresponding to the current feature point in each frame of the multi-frame images;

[0054] Obtain the optimized state data corresponding to each feature point of the target to be tracked, and generate the point cloud data based on the feature point position data and the feature point gray data in the optimized state data.

[0055] Further, generating a 3D Gaussian distribution of the current scene based on the point cloud data and the key frame, and rendering the 3D Gaussian distribution to obtain a first 3D rendered image of the current scene, including:

[0056] Represent each feature point in the point cloud data by the 3D Gaussian distribution;

[0057] Initialize the 3D Gaussian points in the 3D Gaussian distribution, where the initialization is configured to include: project the 3D Gaussian distribution onto a 2D image plane, obtain the rotation matrix R and the scaling matrix S during the projection, and parameterize the covariance matrix ∑ of the 3D Gaussian distribution based on the rotation matrix R and the scaling matrix S;

[0058] Optimize the parameters of the 3D Gaussian distribution based on the key frame and the camera position, the camera speed, and the camera pose corresponding to the key frame, so as to minimize the projection error of the 3D Gaussian distribution;

[0059] Adaptively adjust the number N of 3D Gaussian points in the 3D Gaussian distribution based on the scene complexity metric C, where the number N is configured as:

[0060] N = N 0 ×(1 + α×C)

[0061] In the formula, N 0 is the base number, and α is the adjustment factor;

[0062] Perform differentiable rasterization processing on the 3D Gaussian distribution, and render the processed 3D Gaussian distribution to generate a first 3D rendered image of the current scene.

[0063] The second aspect of the present disclosure provides a 3D scene reconstruction system based on adaptive key frames, including:

[0064] A data acquisition unit, configured to acquire a continuous multi-frame image of the current scene and inertial measurement unit data, where the inertial measurement unit is disposed inside the camera that captures the image and is equipped with a plurality of inertial measurement components;

[0065] A first preprocessing unit, configured to perform target feature extraction and target tracking on each of the multi-frame images respectively to obtain feature point data of the target to be tracked;

[0066] A second preprocessing unit, configured to obtain camera trajectory data and bias data of each of the inertial measurement components based on the inertial measurement unit data;

[0067] A third preprocessing unit, configured to determine key frames based on the scene complexity index and the system resource utilization index of the current scene; and,

[0068] A data optimization unit, configured to perform visual-inertial joint optimization based on all of the feature point data, the camera trajectory data, and the bias data, and generate point cloud data including feature points of the target to be tracked, where the visual-inertial joint optimization is configured to eliminate visual reprojection errors of the point cloud data by correcting the feature point data, the camera trajectory data, and the bias data;

[0069] An image generation unit, configured to generate a 3D Gaussian distribution of the current scene based on the point cloud data and the key frames, and render the 3D Gaussian distribution to obtain a first 3D rendered image of the current scene.

[0070] In the 3D scene reconstruction method provided by the embodiments of the present disclosure, by adopting an adaptive key frame selection strategy, key frames are determined based on the scene complexity index and the system resource utilization index of the current scene, and 3D scene reconstruction is performed based on the key frames, achieving the purpose of improving the adaptability of the method, thereby realizing the goal of automatically adjusting the 3D Gaussian distribution according to the scene complexity and hardware conditions, and being able to operate efficiently in different environments. This adaptability enables excellent rendering results in complex and ever-changing actual application scenarios, thus solving the technical problem that the real-time 3D scene reconstruction effect in the related art is limited due to the inability to adaptively adjust key frames according to the scene complexity and hardware conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the related art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the related art. Obviously, the following drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0072] Figure 1 It is a schematic flowchart of the method provided by the embodiments of the present disclosure;

[0073] Figure 2 It is a software function diagram designed based on the method provided by the embodiments of the present disclosure;

[0074] Figure 3 It is a schematic rendering flowchart provided by the embodiments of the present disclosure;

[0075] Figure 4 Schematic flowchart of the adaptive key frame selection strategy provided by the embodiments of the present disclosure;

[0076] Figure 5 System block diagram provided by the embodiments of the present disclosure.

[0077] Figure 6 Block diagram of the electronic device provided by the embodiments of the present disclosure. Detailed implementation manners

[0078] In order to enable those skilled in the art to better understand the solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present disclosure.

[0079] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so as to describe the embodiments of the present disclosure here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0080] In the present disclosure, the orientation or positional relationship indicated by the terms "upper", "lower", "left", "right", "front", "rear", etc. is based on the orientation or positional relationship shown in the drawings. These terms are mainly used to better describe the present disclosure and its embodiments, and are not used to limit that the indicated devices, elements or components must have a specific orientation, or be constructed and operated in a specific orientation.

[0081] Moreover, in addition to being used to represent an orientation or positional relationship, some of the above terms may also be used to represent other meanings. For example, the term "upper" may also be used to represent a certain attachment relationship or connection relationship in some cases. For those of ordinary skill in the art, the specific meanings of these terms in the present disclosure can be understood according to specific circumstances.

[0082] In addition, the terms "arranged", "provided with", "connected", and "linked" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral structure; it can be a mechanical connection or an electrical connection; it can be a direct connection, or an indirect connection through an intermediate medium, or an internal connection between two devices, components, or parts. For those of ordinary skill in the art, the specific meanings of the above terms in this disclosure can be understood according to specific circumstances.

[0083] The term "3D Gaussian Splatting", abbreviated as 3DGS (the full English name is 3D Gaussian Splatting), is a rasterization technology described by 3D Gaussian distribution for real-time radiation field rendering, which combines high-quality image performance and real-time rendering capabilities. By applying the three core elements of three-dimensional Gaussian distribution, optimizing anisotropic covariance to accurately describe the scene, and developing a fast visibility-aware rendering algorithm, its excellent visual quality and real-time rendering effects have been demonstrated on multiple datasets.

[0084] The term "Inertial Measurement Unit device", abbreviated as IMU (the full English name is Inertial Measurement Unit), is usually composed of components such as accelerometers and gyroscopes. Some IMUs also integrate magnetometers.

[0085] The term "Inertial Measurement Unit data", abbreviated as IMU data, covers various information including accelerometer data and gyroscope data.

[0086] The term "Inertial Measurement Unit bias data", abbreviated as IMU bias, refers to the deviation between the actual value and the theoretical value output by the IMU under specific conditions. This bias is mainly caused by various factors, including manufacturing errors, environmental factors, and wear after long-term use.

[0087] The term "KLT sparse optical flow algorithm" is a classic optical flow estimation method specifically used to track the positions of specific points in an image sequence. The algorithm tracks by selecting appropriate feature points. First, it uses corner detection technology to obtain the sparse feature point set of the previous frame image, and then uses these feature points to match with the subsequent frame image to identify the point set of similar features and infer the optical flow of the sparse point set between the two frame images.

[0088] The term "FAST corner detector" is a feature point detection algorithm widely adopted in the field of computer vision. The full English name of FAST is Features from Accelerated Segment Test. The basic principle of the FAST corner detector is to determine whether a pixel is a corner point by comparing the brightness values of a pixel and its surrounding pixels. Specifically, a circle with a radius of 3 is drawn around the target pixel, and there are 16 pixel points on the circumference. If among these 16 pixel points, there are N consecutive pixels whose brightness values are either brighter or darker than the brightness value of the central pixel plus or minus a threshold, then the central pixel is determined to be a corner point.

[0089] The term "Random Sample Consensus algorithm", abbreviated as RANSAC algorithm (the full English name is RANdom SAmple Consensus), is an iterative algorithm widely used in the fields of computer vision and image processing. By repeatedly randomly selecting data subsets and verifying their consistency, the optimal model parameters are finally determined.

[0090] The term "texture complexity" describes the complexity of texture features in an image. In technical fields such as image processing and video coding and decoding, texture complexity is a core concept, which is related to multiple aspects such as the regularity, variation range, and information volume of image textures. Usually, texture complexity can be quantitatively analyzed through technical means such as the Gray Level Co-occurrence Matrix (GLCM) to extract features such as the mean, variance, and entropy of the texture.

[0091] The term "Simultaneous Localization and Mapping", abbreviated as SLAM (the full English name is Simultaneous Localization and Mapping), is one of the key technologies in the fields of robotics, autonomous driving, and augmented reality. It endows machines with the ability to autonomously navigate in unknown environments and can construct a map of the surrounding environment.

[0092] The term "view synthesis method", abbreviated as NeRF (the full English name is Neural Radiance Fields), is an advanced view synthesis method. The core idea is to use volume rendering technology and simulate the interaction between light and the scene through a deep neural network to achieve high-quality 3D reconstruction and view synthesis.

[0093] The term "Levenberg-Marquardt algorithm", abbreviated as LM algorithm (full name in English is Levenberg-Marquardt), is a widely used nonlinear least squares algorithm. It combines the advantages of the gradient method and the Newton method. When the lambda value in the algorithm is small, the step size is close to the Newton method step size; when the lambda value is large, the step size is close to the step size of the gradient descent method. The LM algorithm is insensitive to the over-parameterization problem, can effectively deal with the redundant parameter problem, and significantly reduces the risk of the cost function falling into a local minimum. In essence, it uses the sum of squared errors as the cost function when solving the nonlinear least squares problem, and finds a set of parameters that minimizes the cost function. In each iteration, the parameter values ​​are updated by calculating the first and second derivatives of the cost function. In order to prevent overfitting, regularization methods are also used to deal with noise and outliers.

[0094] The term "Bag of Words" is a commonly used text representation method in natural language processing and information retrieval. It treats text as a collection of words, counts the number of occurrences of each word, and uses these statistics to construct a feature vector for the text. The core idea of ​​this method is to treat the words in the text as independent features, ignoring the order and grammatical structure between words.

[0095] The term "3D Gaussian Splatting" refers to a technique that converts discrete data points or voxels into a continuous surface or volume. In 3D Gaussian Splatting, each data point or voxel is treated as a Gaussian kernel whose values ​​present a Gaussian distribution in space. By superimposing these Gaussian kernels, a continuous surface or volume can be generated. This superposition process is usually achieved by interpolating each data point or voxel to obtain a smooth and continuous distribution in space.

[0096] The term "GPD radix sort algorithm", abbreviated as NeRF (the full name in English is GPD radix sort), its basic idea is to start from the lowest bit, sort according to the value of the current bit to get a new sequence, and re-sort the new sequence according to the value of the higher bit, and so on until the highest bit, the obtained sequence is an ordered sequence. The following example briefly introduces radix sort.

[0097] The term "pre-integration processing" is a key technology used to reduce the computational burden in visual inertial odometry systems. Its main purpose is to process IMU data to calculate the relative pose transformation between two key frames by assuming that the sensor bias is unchanged. The effects of white noise and bias are considered, and the white noise is stripped off, and the bias is updated through the first-order derivative, so that even if the bias changes, there is no need to recalculate the pre-integration, which improves the optimization efficiency.

[0098] The term "inverse depth parameter", with the full English name of Inverse Depth Parametrization, is the reciprocal of depth. In computer vision, depth usually refers to the distance from the camera's lens or sensor to the surface of an object, that is, the Z-axis distance in the three-dimensional coordinates of a point. And inverse depth is the reciprocal of this distance.

[0099] The term "visual reprojection error model", which is the calculation model of visual reprojection error, is a key concept in computer vision and is mainly used to evaluate and optimize the homography matrix and projection matrix. By minimizing the visual reprojection error, the accuracy in pose estimation and structure from motion can be effectively improved. Visual reprojection error refers to the difference between the observed value and the actual value of a point after projecting a three-dimensional space point onto a two-dimensional plane. Specifically, the reprojection error model is obtained by comparing the pixel coordinates (the observed projection position) with the position obtained by projecting the three-dimensional point according to the currently estimated pose.

[0100] The term "2D feature undistortion processing" refers to calibrating and removing distortion from an image so that the original image can be presented more realistically and accurately.

[0101] The term "pre-integration" uses the continuous measurement data of the IMU to estimate the motion of the camera between two frames of images. This method can reduce the dependence on high-frequency measurements, thereby improving the stability of the algorithm.

[0102] The term "pre-integration residual model" is an important part of visual-inertial odometry (VIO) and is mainly used to improve the positioning accuracy and robustness. By fusing visual and inertial data, the pre-integration residual model can effectively reduce the dependence on high-frequency measurements, thereby improving the overall performance.

[0103] The term "nonlinear optimization" is a mathematical theory and method for studying how to find the optimal solution (usually the minimum value) when the objective function or constraint conditions are nonlinear functions.

[0104] The term "global map loop closure" is a key link in simultaneous localization and mapping, especially significant in the fields of autonomous driving and robot navigation. It involves eliminating the drift of the poses of key frames and landmarks within the loop through optimization algorithms, thereby improving the accuracy and robustness of the system.

[0105] It should be noted that, without conflict, the embodiments and features in the embodiments in this disclosure may be combined with each other. The following will detail this disclosure with reference to the drawings and in combination with the embodiments.

[0106] There are technical problems in related technologies where the real-time 3D scene reconstruction effect is limited because key frames cannot be adaptively adjusted according to the scene complexity and hardware conditions.

[0107] To solve the above technical problems, an embodiment of the present disclosure provides a 3D scene reconstruction method based on adaptive key frames, as Figure 1 shown. The method includes the following steps S11 to S16.

[0108] Step S11: Obtain a series of consecutive frames of images of the current scene and inertial measurement unit data (abbreviated as IMU data), where the inertial measurement unit is disposed inside the camera that captures the images and is equipped with multiple inertial measurement components.

[0109] Step S12: Respectively perform target feature extraction and target tracking on each of the series of frames of images to obtain feature point data of the target to be tracked.

[0110] Step S13: Obtain camera trajectory data and bias data of each inertial measurement component based on the inertial measurement unit data.

[0111] Step S14: Determine key frames based on the scene complexity index and system resource utilization index of the current scene.

[0112] Step S15: Perform visual-inertial joint optimization based on all the feature point data, camera trajectory data, and bias data to generate point cloud data including the feature points of the target to be tracked, where the visual-inertial joint optimization is configured to eliminate the visual reprojection error of the point cloud data by correcting the feature point data, camera trajectory data, and bias data.

[0113] Step S16: Generate a 3D Gaussian distribution of the current scene based on the point cloud data and the key frames, and render the 3D Gaussian distribution to obtain a first 3D rendered image of the current scene.

[0114] Among them, step S11 is used to complete the data acquisition function, and different data acquisition methods can be adopted; step S12 is used to complete the visual feature extraction and tracking function, and different visual feature extraction and tracking algorithms can be adopted; step S13 is used to complete the IMU data processing function, and different IMU data processing algorithms can be adopted; step S14 is used to complete the adaptive key frame selection function, and the key frames are determined based on the scene complexity index and system resource utilization index of the current scene; step S15 is used to complete the visual-inertial joint optimization function, and different visual-inertial joint optimization algorithms can be adopted; step S16 is used to complete the 3D Gaussian representation and rendering function, and different rendering algorithms can be adopted. The embodiments of the present disclosure do not limit any of the above.

[0115] Some embodiments of the present disclosure also provide a 3D scene reconstruction system, a storage medium, and an electronic device corresponding to the above method.

[0116] The 3D scene reconstruction method provided by the embodiments of the present disclosure is applicable to any existing 3D scene reconstruction scenario. For example, in simulating a multi-map scenario in a large-scale environment, applying the 3D scene reconstruction method provided by the embodiments of the present disclosure can maintain efficient operation in different environments compared with the prior art. This adaptability makes the 3D scene reconstruction method perform particularly well in complex and variable actual application scenarios.

[0117] Based on the method of Figure 1 , preferably, the camera trajectory data is configured to include camera position data, camera speed data, and camera attitude data. The inertial measurement component is configured to include an accelerometer and a gyroscope. The feature point data is configured to include feature point position data and feature point grayscale data.

[0118] Based on the method shown in Figure 1 , preferably, the method further includes the following steps S21 - step S23 to complete the function of globally consistent optimization processing of the current scene according to the historical scene.

[0119] Step S21: Identify similar scenes in the historical scene that match the current scene.

[0120] Step S22: Obtain a camera pose calculation model, and perform global consistency optimization on the camera pose calculation model based on the camera trajectory data of the similar scene and the camera trajectory data of the current scene.

[0121] Step S23: Obtain the optimized camera trajectory data of the current scene based on the optimized camera pose calculation model.

[0122] As a preferred implementation, step S21 uses a bag-of-words model to perform scene recognition and quickly detect potential closed loops in the historical scene. This method can efficiently identify similar scenes and provide support for closed-loop detection. The calculation formula for its score is as follows:

[0123] score = ∑tf(w i )·idf(w i ) (1)

[0124] In the formula, tf(w i ) is the term frequency, idf(w i ) is the inverse document frequency, and w i is the feature weight.

[0125] To improve the accuracy of closed-loop detection, the RANSAC algorithm is used for geometric verification to estimate the relative pose and filter out false closed-loop candidates. Its estimation formula is as follows:

[0126] [R,t] = argmin∑||x′i -(Rx i +tt)|| 2 (2) wherein, R is a rotation matrix, tt is a translation vector, and x' i is the three-dimensional coordinate of the matching point, and || || 2 is the squared norm.

[0127] In step S22, the observability of the visual inertial system is enhanced by constructing a pose graph and performing global optimization with four degrees of freedom (x, y, z, yaw), where (x, y, z) are the three-dimensional coordinates of the feature points and yaw is the yaw angle, thereby improving the global consistency. The global consistency optimization is configured to include:

[0128] min{∑||r i,j || 2 +∑ρ(||r i,j || 2 )}(3)

[0129] wherein, r i,j represents the residual of the relative pose constraint, and ρ(·) is a robust kernel function.

[0130] Step S24: Based on the optimized camera trajectory data, perform visual-inertial joint optimization, and then generate a second 3D rendered image of the current scene.

[0131] By combining steps S11 - S16 and steps S21 - S24, the update of global feature points and the reallocation strategy of Gaussian points are realized to ensure the consistency between the current scene image and the optimized pose. This process improves the long-term operation stability of the system by maintaining the global consistency of the system.

[0132] Based on the method shown in Figure 1 , preferably, the method further includes the following steps S31 - S34.

[0133] Step S31: Generate a first sub-map based on the point cloud data of the current scene.

[0134] Step S32: Identify a second scene in the historical scene that forms a global map closed loop with the current scene.

[0135] Step S33: Obtain the point cloud data of the second scene, and generate a second sub-map based on the point cloud data of the second scene.

[0136] Step S34: Merge the first sub-map and the second sub-map to generate a global map.

[0137] Through the combination of steps S11 - S16 and steps S31 - S34, the update of the global map and the re - allocation strategy of Gaussian points are realized to ensure the consistency between the map and the optimized pose. The map merging function is realized, which supports the fusion of multiple sub - maps, can handle the multi - map system in a large - scale environment, and supports advanced application scenarios such as collaborative mapping.

[0138] In some embodiments of the present disclosure, for the above - mentioned step S12, in order to achieve good visual feature extraction and tracking functions, preferably, step S12 further includes the following sub - steps S121 - S125.

[0139] Sub - step S121: For each frame in a multi - frame image as the current frame, perform feature tracking processing on the current frame through the KLT sparse optical flow algorithm, determine the first - type feature points of the target to be tracked, and obtain the feature point data of the first - type feature points, where the feature point data is configured to include feature point position data and feature point gray - scale data.

[0140] Specifically, for each new image, use the KLT sparse optical flow algorithm to track existing features. This algorithm is based on minimizing the following objective function E:

[0141] E=∫∫[I(x,y,t)-I(x + Δx,y + Δy,t + Δt)] 2 dxdy (4)

[0142] In the formula, I(x,y,t) is the gray - scale value of the image at the position (x,y) at time t, and I(x + Δx,y + Δy,t + Δt) is the gray - scale value of the image at the position (x + Δx,y + Δy) at time t + Δt.

[0143] Sub - step S122: Search for corner features of the current frame through the FAST corner detector, determine the second - type feature points of the target to be tracked, and obtain the feature point data of the second - type feature points.

[0144] Specifically, use the FAST corner detector to detect new corner features to maintain 100 - 300 feature points per frame. The judgment criterion is: |I p –I c |>th. Where I p is the intensity of the candidate point, I c is the intensity of the central pixel, and th is the threshold value.

[0145] Sub - step S123: Obtain the minimum pixel interval between adjacent feature points in the feature point set composed of the first - type feature points and the second - type feature points.

[0146] Sub-step S124: Perform distortion removal processing on the 2D features of the target to be tracked for the elements of the feature point set based on the minimum pixel interval, and project the feature point data corresponding to the elements of the processed feature point set onto the unit sphere.

[0147] Specifically, enforce uniform distribution, set the minimum pixel interval between adjacent feature points, perform distortion removal processing on the 2D features, and then project onto the unit sphere:

[0148]

[0149] In the formula, image_area is the obtained area of the image, desired_feature_count is the desired number of feature points, d is the distance value, and d is configured to be equal to the minimum pixel interval used when performing uniform distribution of feature points.

[0150] Sub-step S125: Remove outliers from the projected feature point data through the Random Sample Consensus algorithm (RANSAC algorithm) to obtain the final feature point data of the target to be tracked.

[0151] Specifically, use the RANSAC algorithm to remove outliers, and adopt the fundamental matrix model:

[0152] x′ T Fx = 0 (6)

[0153] In the formula, F is the fundamental matrix, and x and x′ are adjacent feature points.

[0154] In some embodiments of the present disclosure, for step S13, in order to achieve a better IMU data processing function, preferably, step S13 further includes the following sub-steps S131 - sub-step S135.

[0155] Sub-step S131: Obtain camera trajectory data based on inertial measurement unit data.

[0156] Sub-step S132: Obtain the measurement model of the inertial measurement unit, and within the time interval between two consecutive image frames, perform pre-integration processing on the inertial measurement unit data based on this measurement model to obtain the relative displacement, relative velocity, and relative rotation angle of the camera within this time interval.

[0157] Specifically, the measurement model is configured to include:

[0158]

[0159] In the formula, is the actual value of the accelerometer, a t is the measured value of the accelerometer, is the bias data of the accelerometer, is the rotation matrix, g w is the gravitational acceleration, n a is the noise term of the accelerometer, is the actual value of the gyroscope, ω t is the measured value of the gyroscope, is the bias data of the gyroscope, n w is the noise term of the gyroscope.

[0160] The relative displacement, relative velocity and relative rotation angle are respectively configured as:

[0161]

[0162] In the formula, is the relative displacement, is the relative velocity, is the relative rotation angle, is the rotation matrix, Ω is the angular velocity, is the attitude in the local coordinate system, t is the time, t k 、t k+1 are the times corresponding to two consecutive image frames respectively.

[0163] Sub-step S134: Determine the covariance propagation matrix of the measurement model based on the relative displacement, relative velocity and relative rotation angle.

[0164] Specifically, the covariance propagation matrix of the measurement model is configured to include:

[0165]

[0166] In the formula, is the covariance matrix at time t + δt, δt is the time interval, I is the identity matrix, is the covariance matrix at time t, F t is the system matrix, G t is the noise driving matrix, Q t is the noise covariance matrix.

[0167] Among them, is obtained through iterative calculation, describing the uncertainty of state estimation. F t is obtained by taking the partial derivative of the motion equation in the measurement model, describing the linearized model of state transition. G t is obtained by taking the partial derivative of the motion equation in the measurement model with respect to the noise term, describing the influence of system noise on the state variables. Q t is obtained through sensor specifications or experimental calibration, describing the statistical characteristics of system noise.

[0168] Sub-step S135: Optimize the parameters of the measurement model based on the covariance propagation matrix so that the bias data of the inertial measurement components converge to their respective set threshold ranges, and end the optimization.

[0169] Sub-step S136: Obtain the bias data of each inertial measurement component based on the measurement model with optimized parameters.

[0170] As a preferred implementation, the method further includes: identifying that the bias data of any inertial measurement component exceeds the corresponding set threshold, calibrating the relative displacement, relative velocity, and relative rotation angle based on the measurement model, and further calibrating the camera trajectory data.

[0171] In some embodiments of the present disclosure, for step S14, in order to achieve a better adaptive key-frame selection function, preferably, step S14 further includes the following sub-steps S141 - sub-step S145.

[0172] Sub-step S141: For each frame in a multi-frame image as the current frame, obtain the number of feature points N f , the variance of feature point distribution V f , and the texture complexity T f .

[0173] Sub-step S142: Based on the number of feature points N f , the variance of feature point distribution V f , and the texture complexity T f obtain the scene complexity index C of the current frame.

[0174] The scene complexity index C can be configured as:

[0175] C = w 1 ×N f +w 2 ×V f +w 3 ×T f (10)

[0176] In the formula, w 1 , w 2 , w 3 are weight coefficients.

[0177] The above evaluation method can comprehensively consider the geometric and texture features of the scene.

[0178] Sub-step S143: Determine the nearest key frame, and obtain the relative displacement A, overlap degree B, and information entropy H of the current frame relative to the nearest key frame.

[0179] Specifically, the nearest key frame is the key frame with the smallest relative displacement from the current frame.

[0180] The information entropy H is configured as:

[0181] H = -∑p(x i )logp(x i ) (11)

[0182] Wherein, p(x i ) is the probability distribution of the feature points in the image region x i .

[0183] The calculation methods of the relative displacement A and the overlap degree B are common general knowledge in the art, and the embodiments of the present disclosure will not elaborate and limit them herein.

[0184] Sub-step S144: Obtain the system resource utilization rate index R, and dynamically adjust the selection threshold T based on the scene complexity index C and the system resource utilization rate index R.

[0185] Specifically, the selection threshold T is configured as:

[0186] T = T 0 ×(1 + α×C)×(1 - β×R) (12)

[0187] Wherein, T 0 is the base threshold, α is the adjustment coefficient, and β is the correction coefficient. This mechanism can lower the selection criteria in complex scenes, raise the criteria in simple scenes, and consider the system load at the same time.

[0188] Sub-step S145: Determine whether the current frame is a key frame based on the information entropy H, the selection threshold T, the relative displacement A, and the overlap degree B.

[0189] As a preferred implementation manner, the key frame is configured to simultaneously satisfy: the information entropy H exceeds the selection threshold T, and, the relative displacement A exceeds the preset maximum allowable value, and the overlap degree B is lower than the preset overlap threshold.

[0190] In some embodiments of the present disclosure, for step S15, in order to achieve a better visual-inertial joint optimization function, step S15 may further include the following sub-steps S151 - sub-step S15.

[0191] Sub-step S151: For each feature point of the target to be tracked as the current feature point, obtain the corresponding state data of the current feature point in each frame of the multi-frame image, wherein the state data is configured to include feature point position data, feature point gray data, camera position data, camera speed data, camera attitude data, and bias data of each inertial measurement component.

[0192] Sub-step S152: Obtain the inverse depth parameter corresponding to the current feature point in each frame of the multi-frame image, and generate the state vector X of the current feature point based on the state data and the inverse depth parameter.

[0193] Specifically, the state vector X can be configured as follows:

[0194]

[0195] wherein, x i is the state data of the i-th frame, and n is the number of frames of multiple images. is the configuration state vector of the camera, which is used to describe the current state of the camera, including parameters such as the position, speed, and attitude of the camera. λ 0 , λ 1 , …, λ m are the inverse depth parameters of the current feature points, which are used to describe the depth information of the feature points in the camera coordinate system. Through the inverse depth parameters, optimization and calculation can be more efficiently performed during the projection process of visual feature points.

[0196] A complete state vector including the camera pose (camera position data, camera speed data, camera attitude data), IMU bias, and feature point positions, etc., is constructed through sub-step S152 for the joint optimization of the visual-inertial system. This state vector can unify all variables in the system, thereby achieving accurate estimation of the camera trajectory and feature point positions.

[0197] Sub-step S153: Input the state vector X into the visual reprojection error model to obtain the visual reprojection error.

[0198] Specifically, the visual reprojection error in the visual reprojection error model is configured as follows:

[0199]

[0200] wherein, π() is the projection function, is the camera pose transformation matrix, p l is the three-dimensional coordinate of the current feature point, is the two-dimensional coordinate of the current feature point projected on the 2D image plane.

[0201] The camera pose transformation matrix fuses the pre-integration information of the IMU and the visual information. By establishing the relationship between the pre-integration measurement and the state vector, tight-coupling optimization of the visual-inertial data is achieved.

[0202] By defining the visual reprojection error, the difference between the projection of the feature point on the image and the actual observation is measured. The visual reprojection error is the core constraint in visual simultaneous localization and mapping (SLAM), which is used to correct the camera pose and feature point positions.

[0203] Sub-step S154: Input the state vector X into the pre-integration residual model to obtain the pre-integration residual.

[0204] Specifically, the pre-integration residual in the pre-integration residual model is configured as:

[0205]

[0206] wherein, δα is the error term of the camera position, δβ is the error term of the camera velocity, δθ is the error term of the camera attitude, δb a is the bias error of the accelerometer, and δb w is the bias error of the gyroscope, and is the parameter vector to be estimated.

[0207] To make full use of the IMU data, the pre-integration residual is defined, and this residual data measures the consistency between the IMU pre-integration result and the state estimation. The introduction of this residual data helps to improve the smoothness of the trajectory estimation and the accuracy of the overall system.

[0208] Sub-step S155: Based on the visual reprojection error and the pre-integration residual, perform non-linear optimization on the state vector X to obtain the optimized state data corresponding to the current feature point in each frame of multiple frames of images.

[0209] Specifically, the above non-linear optimization is configured as:

[0210]

[0211] wherein, r p is the prior residual, H p is the Jacobian matrix of the prior information, P is the inverse of the covariance propagation matrix, ρ(·) is the robust kernel function, and is the Euclidean norm.

[0212] Sub-step S155 combines various constraints of vision and IMU by constructing a non-linear optimization problem, and the optimization objective is to minimize the weighted sum of all error terms. This optimization process not only considers the visual reprojection error and the pre-integration residual, but also includes the prior information, so as to achieve the global optimization of the camera pose, the IMU state, and the feature point position.

[0213] As a preferred embodiment, to effectively solve the constructed non-linear optimization problem, the Levenberg-Marquardt algorithm can be used for solution. This algorithm combines the fast convergence characteristics of the Gauss-Newton method and the stability of the gradient descent method, and can handle the non-linear problems in the SLAM system. The linearized formula of the optimization problem is:

[0214] (J T J + λI)δx = -J T r (17)

[0215] In the formula, J is the Jacobian matrix, r is the residual vector, λ is the damping factor, and δx is the increment of the optimization variable. By iteratively solving this equation, the camera pose, the IMU bias, and the feature point positions can be gradually optimized.

[0216] Sub-step S156: Obtain the optimized state data corresponding to each feature point of the target to be tracked, and generate point cloud data based on the feature point position data and the feature point gray-scale data in the optimized state data.

[0217] In some embodiments of the present disclosure, for step S16, in order to achieve a better 3D Gaussian distribution and rendering effect, preferably, step S16 further includes the following sub-steps S161 - sub-step S165.

[0218] Sub-step S161: Represent each feature point in the point cloud data by a 3D Gaussian distribution.

[0219] Specifically, a 3D Gaussian distribution is used to represent each point in the scene. The Gaussian distribution of each point is described by the mean value and the covariance matrix, which can effectively capture the spatial uncertainty of the point. By using the 3D Gaussian distribution, the position of the point and its uncertainty can be accurately described in the scene reconstruction. The definition of the 3D Gaussian distribution is as follows:

[0220]

[0221] In the formula, μ is the mean value of the Gaussian distribution, and ∑ is the covariance matrix, which is used to describe the spatial distribution characteristics of the point.

[0222] Sub-step S162: Initialize the 3D Gaussian points in the 3D Gaussian distribution, where the initialization is configured to include: projecting the 3D Gaussian distribution onto the 2D image plane, obtaining the rotation matrix R, the scaling matrix S during the projection process, and parameterizing the covariance matrix ∑ of the 3D Gaussian distribution based on the rotation matrix R and the scaling matrix S.

[0223] Specifically, the relationship between the rotation matrix R, the scaling matrix S, and the covariance matrix ∑ is established through the following formula:

[0224] ∑ = RSS T R T (19)

[0225] Projecting the 3D Gaussian distribution onto the 2D image plane can achieve an effective analysis of the three-dimensional space. During the projection process, the internal and external parameters of the camera and the characteristics of the Gaussian distribution are considered, so that the information in the three-dimensional space can be accurately transformed into the two-dimensional image. The projection expression of the 3D Gaussian distribution is as follows:

[0226] ∑′ = JW∑W TJ T (20)

[0227] In the formula, J is the Jacobian matrix of the projection transformation, and W is the transformation matrix from the current scene to the camera.

[0228] To ensure the positive definiteness of the covariance matrix and the stability of the optimization process, a scaling matrix S and a rotation matrix r are used to parameterize the covariance matrix ∑, as shown in formula (19). This parameterization method not only ensures the positive definiteness of the covariance matrix but also provides an intuitive geometric interpretation.

[0229] Sub-step S163: Optimize the parameters of the 3D Gaussian distribution based on the key frames and the corresponding camera positions, camera speeds, and camera poses, so as to minimize the projection error of the 3D Gaussian distribution.

[0230] Specifically, by optimizing the parameters of the Gaussian distribution, including the position parameter μ, the scaling parameter s, and the rotation parameter q, its projection error is minimized. The parameters of the Gaussian distribution are adjusted from multiple perspectives to make its projection more consistent with the observations. The optimization objective is as follows:

[0231]

[0232] In the formula, π is the projection function, z is the actual observation value, and G(μ, s, q) is the Gaussian distribution. The specific form of the Gaussian distribution will not be elaborated. Finally, the parameters of the optimized 3D Gaussian distribution are obtained, including the position parameter μ, the scaling parameter s, and the rotation parameter q.

[0233] Sub-step S164: Adaptively adjust the number N of 3D Gaussian points in the 3D Gaussian distribution based on the scene complexity index C, where the number N is configured as:

[0234] N = N 0 ×(1 + α × C) (22)

[0235] In the formula, N 0 is the base number, and α is the adjustment factor.

[0236] To improve the accuracy in complex areas and reduce the computational amount in simple areas, a mechanism for adaptively adjusting the number of Gaussian points according to the scene complexity is proposed. The number of Gaussian points is adjusted according to the scene complexity to balance the reconstruction quality and the computational efficiency.

[0237] Sub-step S165: Perform differentiable rasterization processing on the 3D Gaussian distribution, and render the processed 3D Gaussian distribution to generate the first 3D rendered image of the current scene.

[0238] As a preferred implementation manner, sub-step S165 may further include the following sub-steps 165.1 - sub-step S165.5.

[0239] Sub-step 165.1: Project Gaussian points onto the 2D image plane using 3D Gaussian sputtering technology. This method fully considers the spatial characteristics of the Gaussian distribution and can generate smooth rendering results, ensuring the accurate projection of three-dimensional scene information onto the two-dimensional image. The projection formula is as follows:

[0240] p 2D = K[R|t]·p 3D (23)

[0241] where K is the camera intrinsic matrix, [R|t] is the camera extrinsic matrix, p 2D is the 2D coordinate of the feature point, and p 3D is the 3D coordinate of the feature point.

[0242] Sub-step 165.2: Consider the deformation generated during the 3D Gaussian projection by calculating the covariance of the 2D Gaussian distribution, ensuring that the rendering result can accurately reflect the performance of the 3D structure on the 2D plane. The calculation formula for the 2D Gaussian covariance is:

[0243] Σ 2D = J·∑ 3D ·J T (24)

[0244] where J is the Jacobian matrix of the projection, and ∑ 3D is the covariance matrix of the 3D Gaussian projection.

[0245] Sub-step 165.3: Rasterize the Gaussian points using the Elliptical Weighted Average (EWA) filtering technique. EWA filtering can produce high-quality anti-aliasing effects, thereby improving the visual quality of the image. The specific filtering formula is as follows:

[0246]

[0247] where r is the pixel radius, ∑ 2D is the 2D covariance matrix, w(x) is the filtered pixel, and x is the original pixel.

[0248] Sub-step 165.4: Use the α-blending algorithm based on depth order and opacity, which can correctly handle the semi-transparent effect and generate a realistic rendering image C. Its formula is:

[0249] C = ∑c i ·α i ·∏(1 - α j ), j < i (26)

[0250] where c i is the color of the Gaussian point, and α i , α j are the opacities.

[0251] Sub-step 165.5: Use spherical harmonics (SH) to capture the view-dependent lighting effect for the directional appearance. The use of spherical harmonics can efficiently simulate complex lighting changes and improve the quality of the rendered image. The formula is as follows:

[0252] L(ω) = ∑c i ·Y i (ω) (27)

[0253] where c i is the SH coefficient, Y i (ω) is the spherical harmonic basis function, L(ω) is the irradiance, and ω is the direction vector.

[0254] As a preferred implementation, in order to better render the 3D Gaussian distribution, GPU-accelerated real-time rendering can be adopted. Refer to the following sub-steps S165.6 - sub-step S165.9.

[0255] Sub-step S165.6: Implement a tile-based parallel rendering strategy, divide the image into multiple small blocks, and perform parallel processing on the GPU. This method makes full use of the parallel computing power of the GPU, significantly improving the rendering speed and is suitable for real-time rendering scenarios. In order to ensure the correct rendering order of the Gaussian points, the GPD radix sorting algorithm is used to sort the depths of the Gaussian points. This sorting algorithm has high performance and can ensure the correct depth order processing while maintaining the real-time rendering performance.

[0256] Sub-step S165.7: Ensure the correctness of transparency processing through a forward-backward traversal α-blending strategy. This blending strategy can handle complex occlusion relationships, thereby improving the rendering quality. The formula is as follows:

[0257] C = ∑c i ·α i ·∏(1 - α j ), j < i (28)

[0258] In the formula, c i is the color of the Gaussian point, and α i , α j are the transparencies of the Gaussian points.

[0259] Sub-step S165.8: Through an adaptive mechanism, dynamically adjust the rendering resolution and the number of Gaussian points according to the hardware conditions to balance the image quality and the frame rate. Under different hardware conditions, a smooth user experience can be maintained. The adjustment formula is as follows:

[0260]

[0261] In the formula, R is the rendering resolution, R0 is the initial rendering resolution, N is the number of Gaussian points, N 0 is the initial number of Gaussian points, f r is the current frame rate, f t is the target frame rate.

[0262] Sub-step S165.9: To optimize the Gaussian parameters, GPU-accelerated gradient calculation significantly improves the efficiency of parameter optimization. The specific gradient calculation formula is as follows:

[0263]

[0264] In the formula, L is the rendering loss, θ is the Gaussian parameter, p i is the pixel value.

[0265] Figure 2 is the software function diagram designed according to the method provided by the present disclosure. In this software, input ports for camera images and IMU data are set, and image processing buttons for restoring depth maps, feature extraction, and tracking are set. After receiving the initialization instruction, the camera pose is predicted through IMU data, and feature points are triangulated. The sliding window is optimized based on the predicted camera pose and feature point triangulation data, and finally, camera pose data, key frames, and point clouds are output.

[0266] The rendering process of this software is as Figure 3 shown, and the adaptive key frame selection strategy is as Figure 4 shown. The rendering process corresponds to the above sub-step S165, and the adaptive key frame selection strategy corresponds to the above step S14, which will not be elaborated here.

[0267] From the above description, it can be seen that the present disclosure achieves the following technical effects:

[0268] 1. Provides a real-time 3D scene reconstruction method that can adaptively adjust key frames, optimize resource allocation, balance rendering quality and real-time performance, and has good robustness.

[0269] 2. Improves the adaptability of key frame selection and the adaptability of the method to the environment. Designs a key frame selection strategy that can dynamically adjust according to scene complexity and system resource status. Through the adaptive key frame selection strategy and dynamic Gaussian point adjustment mechanism, the system can automatically adjust according to scene complexity and hardware conditions and maintain efficient operation in different environments. This adaptability enables the system to perform well in complex and changing actual application scenarios.

[0270] 3. Optimizes the calculation resource allocation. The calculation of resources during the rendering process is fed back to the key frames, and the adaptive key frame strategy improves the calculation efficiency and memory utilization rate of the system.

[0271] 4. Improve the rendering quality and real-time performance. By decoupling the pose and key frames and combining with an optimized key frame strategy, the real-time performance of the system is improved while ensuring the rendering quality.

[0272] 5. Enhance the system robustness. Through technologies such as visual-inertial joint optimization and loop detection, the system has strong anti-interference ability and long-term stability, and can operate reliably in various complex environments.

[0273] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0274] The embodiments of the present disclosure also provide a 3D scene reconstruction system for implementing the above method embodiments, as Figure 5 shown. The 3D scene reconstruction system 20 includes a data acquisition unit 201, a first preprocessing unit 202, a second preprocessing unit 203, a third preprocessing unit 204, a data optimization unit 205, and an image generation unit 206.

[0275] The data acquisition unit 201 is configured to acquire a continuous multi-frame image of the current scene and inertial measurement unit data. Among them, the inertial measurement unit is arranged inside the camera that captures the image and is equipped with multiple inertial measurement components.

[0276] The first preprocessing unit 202 is configured to respectively perform target feature extraction and target tracking on each frame of the multi-frame images to obtain the feature point data of the target to be tracked.

[0277] The second preprocessing unit 203 is configured to obtain the camera trajectory data and the bias data of each inertial measurement component based on the inertial measurement unit data.

[0278] The third preprocessing unit 204 is configured to determine key frames based on the scene complexity index and the system resource utilization index of the current scene.

[0279] The data optimization unit 205 is configured to perform visual-inertial joint optimization based on all the feature point data, camera trajectory data, and bias data and generate point cloud data including the feature points of the target to be tracked. Among them, the visual-inertial joint optimization is configured to eliminate the visual reprojection error of the point cloud data by correcting the feature point data, camera trajectory data, and bias data.

[0280] The image generation unit 206 is configured to generate a 3D Gaussian distribution of the current scene based on the point cloud data and the key frames, and render the 3D Gaussian distribution to obtain the first 3D rendered image of the current scene.

[0281] The specific ways of the operations performed by each unit in the above device embodiments have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0282] Embodiments of the present disclosure also provide an electronic device, as Figure 6 shown, the electronic device includes one or more processors 31 and a memory 32. Figure 6 Here, one processor 31 is taken as an example.

[0283] The controller may further include: an input device 33 and an output device 34.

[0284] The processor 31, the memory 32, the input device 33, and the output device 34 may be connected by a bus or other means. Figure 6 Here, being connected by a bus is taken as an example.

[0285] The processor 31 may be a central processing unit (Central Processing Unit, abbreviated as CPU), and the processor 31 may also be other general-purpose processors, digital signal processors (Digital Signal Processor, abbreviated as DSP), application specific integrated circuits (Application Specific Integrated Circuit, abbreviated as ASIC), field programmable gate arrays (Field-Programmable Gate Array, abbreviated as FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. chips, or a combination of the above various types of chips. The general-purpose processor may be a microprocessor or any conventional processor.

[0286] The memory 32, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the control method in the embodiments of the present disclosure. The processor 31 executes various functional applications and data processing of the server by running the non-transitory software programs, instructions, and modules stored in the memory 32, that is, implements the 3D scene reconstruction method in the above method embodiments.

[0287] The memory 32 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function. The data storage area may store data created according to the use of the processing device operating on the server, etc. In addition, the memory 32 may include a high-speed random access memory and may also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 32 may optionally include a memory remotely provided with respect to the processor 31, and these remote memories may be connected to the network connection device through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0288] The input device 33 may receive input digital or character information and generate key signal inputs related to user settings and function controls of the processing device of the server. The output device 34 may include a display device such as a display screen.

[0289] One or more modules are stored in the memory 32 and, when executed by one or more processors 31, perform the 3D scene reconstruction method as Figure 1 shown.

[0290] Those skilled in the art can understand that to implement all or part of the processes in the above method embodiments, it can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium, and when the program is executed, it may include the processes of the above method embodiments. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory (FM), a hard disk drive (HDD), or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above types of memories.

[0291] Although the embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A 3D scene reconstruction method based on adaptive keyframes, characterized in that: include: Acquire a continuous multi-frame image of the current scene and inertial measurement unit data, wherein the inertial measurement unit is arranged inside a camera that takes the image and is equipped with a plurality of inertial measurement components; Extracting target features and tracking the target are performed on each frame of the multiple frames of images to obtain feature point data of the target to be tracked; Acquire camera trajectory data and bias data of each of the inertial measurement components based on the inertial measurement unit data; Determining a key frame based on a scene complexity index and a system resource utilization index of the current scene; Based on all the feature point data, the camera trajectory data, and the bias data, visual-inertial joint optimization is performed to generate point cloud data containing each feature point of the target to be tracked, wherein the visual-inertial joint optimization is configured to eliminate the visual reprojection error of the point cloud data by correcting the feature point data, the camera trajectory data, and the bias data; and A 3D Gaussian distribution of the current scene is generated based on the point cloud data and the key frame, and the 3D Gaussian distribution is rendered to obtain a first 3D rendered image of the current scene.

2. The method according to claim 1, characterized in that Also includes: Identifying similar scenes in historical scenes that match the current scene; Acquire the camera pose calculation model, and perform global consistency optimization on the camera pose calculation model based on the camera trajectory data of the similar scene and the camera trajectory data of the current scene; Acquire optimized camera trajectory data of the current scene based on the optimized camera pose calculation model; as well as, The visual-inertial joint optimization is performed based on the optimized camera trajectory data to generate a second 3D rendered image of the current scene.

3. The method according to claim 1 or 2, characterized in that: Also includes: Generate a first sub-map based on the point cloud data of the current scene; Identifying a second scene in the historical scene that forms a global map closed loop with the current scene; Acquire point cloud data of the second scene, and generate a second sub-map based on the point cloud data of the second scene; as well as, The first sub-map and the second sub-map are merged to generate a global map.

4. The method according to claim 1 or 2, characterized in that: The determining of the key frame based on the scene complexity index and the system resource utilization index of the current scene includes: For each frame in the multiple frames of the image as the current frame, obtain the number of feature points N of the current frame f , feature point distribution variance V f , texture complexity T f ; Based on the number of feature points N f , the feature point distribution variance V f , the texture complexity T f Obtaining a scene complexity index C of the current frame; Determine the nearest key frame, and obtain the relative displacement A, overlap degree B and information entropy H of the current frame relative to the nearest key frame; Obtaining a system resource utilization index R, and dynamically adjusting a selection threshold T based on the scenario complexity index C and the system resource utilization index R; and, Based on the information entropy H, the selection threshold T, the relative displacement A, and the overlap degree B, it is determined whether the current frame is the key frame.

5. The method according to claim 4, characterized in that The scene complexity index C, the information entropy H, and the selection threshold T are respectively configured as: C=w1×N f +w2×V f +w3×T f H=-∑p(x i )logp(x i ) T=T0×(1+α×C)×(1-β×R) In the formula, w1, w2, w3 are weight coefficients, p(x i ) is the feature point in the image area x i The probability distribution in , T0 is the basic threshold, α is the adjustment coefficient, β is the correction coefficient; and, The key frame is configured to simultaneously satisfy: the information entropy H exceeds the selection threshold T, and the relative displacement A exceeds a preset maximum allowable value, and the overlap degree B is lower than a preset overlap threshold.

6. The method according to claim 5, characterized in that The step of performing target feature extraction and target tracking on each frame of the multiple frames of images to obtain feature point data of the target to be tracked includes: For each frame of the multiple frames of images as a current frame, perform feature tracking processing on the current frame by using a KLT sparse optical flow algorithm to determine a first type of feature point of the target to be tracked, and obtain feature point data of the first type of feature point, wherein the feature point data is configured to include feature point position data and feature point grayscale data; Performing a corner point feature search on the current frame by using a FAST corner point detector to determine a second type of feature point of the target to be tracked, and acquiring feature point data of the second type of feature point; Obtaining a minimum pixel interval between adjacent feature points in a feature point set consisting of the first type of feature points and the second type of feature points; Based on the minimum pixel interval, the elements of the feature point set are subjected to 2D feature dedistortion processing of the target to be tracked, and feature point data corresponding to the elements of the feature point set after the processing are projected onto a unit sphere; and The projected feature point data is subjected to outlier elimination through a random sampling consistency algorithm to obtain final feature point data of the target to be tracked.

7. The method according to claim 1 or 2, characterized in that: The camera trajectory data is configured to include camera position data, camera speed data, and camera attitude data, and the inertial measurement unit is configured to include an accelerometer and a gyroscope; and the camera trajectory data and the bias data of each inertial measurement unit are obtained based on the inertial measurement unit data, including: acquiring camera trajectory data based on the inertial measurement unit data; and, Acquire a measurement model of the inertial measurement unit, and perform pre-integration processing on the inertial measurement unit data based on the measurement model within a time interval between two consecutive image frames to obtain a relative displacement, relative speed, and relative rotation angle of the camera within the time interval; Determine a covariance propagation matrix of the measurement model based on the relative displacement, relative velocity and relative rotation angle; Optimizing the parameters of the measurement model based on the covariance propagation matrix so that the bias data of the inertial measurement assembly converge to respective set threshold ranges, and ending the optimization; The bias data of each of the inertial measurement components is obtained based on the measurement model after parameter optimization.

8. The method according to claim 7, characterized in that Also includes: When it is identified that the bias data of any of the inertial measurement components exceeds the corresponding set threshold, the relative displacement, the relative speed and the relative rotation angle are calibrated based on the measurement model, thereby calibrating the camera trajectory data; and, The performing visual-inertial joint optimization based on all the feature point data, the camera trajectory data, and the bias data and generating point cloud data containing each feature point of the target to be tracked includes: For each feature point of the target to be tracked as a current feature point, obtaining state data corresponding to the current feature point in each frame of the multiple frames of images, wherein the state data is configured to include the feature point position data, the feature point grayscale data, the camera position data, the camera speed data, the camera attitude data, and the bias data of each of the inertial measurement components; Acquire an inverse depth parameter corresponding to the current feature point in each frame of the multiple frames of images, and generate a state vector X of the current feature point based on the state data and the inverse depth parameter; Inputting the state vector X into a visual reprojection error model to obtain a visual reprojection error; Inputting the state vector X into a pre-integrated residual model to obtain a pre-integrated residual; Based on the visual reprojection error and the pre-integrated residual, the state vector X is nonlinearly optimized to obtain the optimized state data corresponding to the current feature point in each frame of the multiple frames of images; and The optimized state data corresponding to each feature point of the target to be tracked is obtained, and the point cloud data is generated based on the feature point position data and the feature point grayscale data in the optimized state data.

9. The method according to any one of claims 1 to 8, characterized in that: The step of generating a 3D Gaussian distribution of the current scene based on the point cloud data and the key frame, and rendering the 3D Gaussian distribution to obtain a first 3D rendered image of the current scene includes: Representing each of the feature points in the point cloud data by the 3D Gaussian distribution; Initializing the 3D Gaussian points in the 3D Gaussian distribution, wherein the initialization is configured to include: projecting the 3D Gaussian distribution onto a 2D image plane, acquiring a rotation matrix R and a scaling matrix S in the projection process, and parameterizing a covariance matrix ∑ of the 3D Gaussian distribution based on the rotation matrix R and the scaling matrix S; Optimizing the parameters of the 3D Gaussian distribution based on the key frame and the camera position, the camera speed, and the camera posture corresponding to the key frame so as to minimize the projection error of the 3D Gaussian distribution; Adaptively adjusting the number N of the 3D Gaussian points in the 3D Gaussian distribution based on the scene complexity index C, wherein the number N is configured as: N=N0×(1+α×C) In the formula, N0 is the basic quantity, and α is the adjustment factor; The 3D Gaussian distribution is subjected to differentiable rasterization processing, and the processed 3D Gaussian distribution is rendered to generate a first 3D rendered image of the current scene.

10. A 3D scene reconstruction system based on adaptive keyframes, characterized in that: include: A data acquisition unit configured to acquire a plurality of consecutive frames of images of a current scene and inertial measurement unit data, wherein the inertial measurement unit is disposed inside a camera that captures the images and is equipped with a plurality of inertial measurement components; A first preprocessing unit is configured to perform target feature extraction and target tracking on each frame of the multiple frames of images to obtain feature point data of the target to be tracked; A second pre-processing unit is configured to acquire camera trajectory data and bias data of each of the inertial measurement components based on the inertial measurement unit data; A third preprocessing unit is configured to determine a key frame based on a scene complexity index and a system resource utilization index of the current scene; a data optimization unit, configured to perform visual-inertial joint optimization based on all the feature point data, the camera trajectory data, and the bias data and generate point cloud data containing each feature point of the target to be tracked, wherein the visual-inertial joint optimization is configured to eliminate the visual reprojection error of the point cloud data by correcting the feature point data, the camera trajectory data, and the bias data; and The image generation unit is configured to generate a 3D Gaussian distribution of the current scene based on the point cloud data and the key frame, and render the 3D Gaussian distribution to obtain a first 3D rendered image of the current scene.

Citation Information

Patent Citations

  • Dense visual scene reconstruction method and system

    CN119478277A

  • Systems and methods for three-dimensional body part modelling

    US20240320392A1

Cited By

  • Large-scene three-dimensional reconstruction method based on three-dimensional Gaussian sputtering

    CN120472121A

  • AR glasses dangerous state dynamic labeling method and system based on multi-source perception

    CN120543957A

  • Automatic scene calibration method based on machine vision

    CN120635605A

  • Dynamic scene reconstruction method using Kalman filter to guide Gaussian rendering

    CN121074215A

  • Method and system for optimizing and accelerating visual inertial navigation system of unmanned aerial vehicle based on heterogeneity

    CN121430604A