Spatial real-time reconstruction method based on deep learning

Through a deep learning-based real-time spatial reconstruction method, combined with optical flow field analysis, Poisson equation depth compensation, conditional random field optimization and incremental map update, the low robustness and accuracy of spatial reconstruction in dynamic scenarios are solved, and the consistent spatial reconstruction effect of high precision, geometric continuity and real-time is achieved.

CN119991958AInactive Publication Date: 2025-05-13MIRROR VISION (ZHEJIANG) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510086704.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing spatial reconstruction methods show low robustness and accuracy in dynamic scenarios. The detection accuracy of dynamic objects is affected by light changes, occlusion and reflections. The depth recovery results lack geometric continuity and boundary consistency, resulting in the impact of model accuracy and coherence.

Method used

The real-time spatial reconstruction method based on deep learning is adopted to achieve accurate segmentation and background recovery of dynamic objects through the combination of optical flow field analysis and depth data; depth compensation is used to ensure the geometric continuity of dynamic regions; global consistency optimization is carried out through conditional random field and graph model to enhance the accuracy and consistency of label allocation; incremental map update method is used to adjust the model content in real time to maintain geometric accuracy and stability.

Benefits of technology

It significantly improves the accuracy of dynamic object detection in dynamic environments, realizes the geometric continuity of dynamic backgrounds and the accuracy of depth information, enhances the spatial and temporal consistency of label segmentation results, and improves the stability and real-timeness of map updates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991958A_ABST
    Figure CN119991958A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of three-dimensional reconstruction, and discloses a space real-time reconstruction method based on deep learning, which comprises the following steps of: 1, preparing input data, namely acquiring an input image sequence, corresponding depth data and camera pose information, acquiring a camera internal reference matrix for conversion from image coordinates to camera coordinates, and acquiring the image sequence; distortion correction and denoising processing are carried out on the input image sequence and depth data, and the processed image sequence and depth data are provided for the step 2 for dynamic object detection and analysis; and step 2, dynamic object detection: calculating time change information and space gradient information of pixel points based on the image sequence processed in the step 1 and the depth data. Through a dynamic object detection method based on optical flow constraint, motion features of a dynamic area are analyzed in combination with pixel point time change and depth data, accurate segmentation and positioning of a dynamic object are achieved, and the effect of remarkably improving the dynamic object detection precision in a dynamic environment is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of three-dimensional reconstruction technology, and specifically to a real-time spatial reconstruction method based on deep learning. Background Art

[0002] With the rapid development of computer vision and artificial intelligence technology, real-time spatial reconstruction technology has been widely used in fields such as autonomous driving, augmented reality, virtual reality, and robot navigation. The core of real-time spatial reconstruction is to model three-dimensional scenes with high precision to achieve environmental perception and dynamic interaction. However, there are many technical challenges in spatial reconstruction in dynamic environments.

[0003] Existing spatial reconstruction methods are mostly based on the traditional simultaneous positioning and mapping framework, relying on static scene assumptions and geometric constraints for environment modeling. However, they show low robustness and accuracy when dealing with dynamic scenes, including the following technical issues:

[0004] Traditional methods mainly rely on optical flow analysis and geometric consistency to detect dynamic objects, but the motion patterns of dynamic objects are complex, and the detection accuracy is easily affected by lighting changes, occlusion and reflection factors, resulting in blurred segmentation boundaries and misjudgment, which affects the accuracy of spatial reconstruction.

[0005] In dynamic scenes, moving objects can cause background occlusion and information loss. Existing restoration methods mostly use simple interpolation and texture matching-based methods to fill in the missing information, which can easily lead to a lack of geometric continuity and boundary consistency in the depth recovery results, thus affecting the overall accuracy and coherence of the reconstructed model.

[0006] There is a fuzzy transition area between the boundaries of static background and dynamic objects in dynamic scenes. The existing deep learning-based segmentation model cannot ensure label consistency and spatial continuity in time series, resulting in label drift and segmentation instability, which in turn affects the accuracy and spatial consistency of the model optimization results.

[0007] Existing methods mainly rely on global optimization and dense reconstruction models during map updates. When faced with large-scale scenes, the storage and computational complexity increases significantly and cannot meet the needs of real-time reconstruction. At the same time, the lack of dynamic management of dynamic object motion and background compensation can easily lead to geometric rupture and drift in modeling results, affecting the stability of real-time mapping.

[0008] Therefore, those skilled in the art provide a deep learning-based real-time spatial reconstruction method to solve the above-mentioned problems. Summary of the invention

[0009] In view of the shortcomings of the prior art, the present invention provides a deep learning-based real-time spatial reconstruction method to solve the problems raised in the above background technology.

[0010] To achieve the above objectives, the present invention is implemented through the following technical solutions: a spatial real-time reconstruction method based on deep learning, comprising:

[0011] Step 1: prepare input data, including obtaining input image sequence and corresponding depth data and camera pose information, obtaining camera intrinsic parameter matrix for transforming image coordinates to camera coordinates, and performing distortion correction and denoising on input image sequence and depth data, and providing the processed image sequence and depth data to step 2 for dynamic object detection and analysis;

[0012] Step 2, dynamic object detection, including calculating the time change information and spatial gradient information of the pixel points based on the image sequence and depth data processed in step 1, extracting the optical flow field by analyzing the displacement change of the pixel points between time frames, analyzing the motion characteristics of the dynamic object area in combination with the depth data, generating a dynamic segmentation mask by setting a threshold to distinguish the dynamic area from the static background area, and passing the generated dynamic segmentation mask and the position information of the dynamic area as input to step 3 for dynamic background recovery;

[0013] Step 3, dynamic background recovery, including determining the dynamic area and the boundary area based on the dynamic segmentation mask generated in step 2 and the depth data provided in step 1, analyzing the boundary pixel information of the dynamic area in the dynamic segmentation mask, constraining the depth compensation of the dynamic area according to the depth value of the boundary area, guiding the recovery of the internal depth value of the dynamic area in combination with the continuity of the boundary depth information, generating a depth map after dynamic background recovery, and providing the restored depth map and the dynamic segmentation mask to step 4 for global consistency modeling and optimization;

[0014] Step 4: Global consistency modeling, including building a graph model of the pixel feature space based on the depth map after dynamic background restoration in step 3 and the dynamic segmentation mask generated in step 2, characterizing the feature differences between dynamic and static regions through edge weights, optimizing the global consistency of dynamic and static regions through a label propagation mechanism, updating pixel and voxel labels, and using the optimized labels and pixel features to describe the spatial distribution of the static background and dynamic foreground of the scene. The optimization results and the depth map after background restoration are provided to step 5 for building a dense 3D model.

[0015] Step 5: dense 3D model construction and update, including generating a dense 3D model based on the pixel features and label information optimized in step 4 and the depth map after dynamic background restoration in step 3, dynamically adjusting the model content frame by frame through an incremental map update method, performing depth compensation on the occluded area in combination with the background restoration result in step 3, and using a spatial division structure to manage the storage and spatial partition calculation of the dense model, and finally outputting the dense 3D model as a complete reconstruction result of the static background, and performing unified coordinate correction on the dense model according to the posture information provided in step 1.

[0016] Preferably, the calculation of the optical flow field in step 2 is based on the following optical flow constraint equation:

[0017] I x u+I y v+I t =0,

[0018] Among them, I x ,I y is the spatial gradient of the image in the horizontal and vertical directions, I t is the temporal gradient of the image, u and v represent the horizontal and vertical displacement of pixels in the optical flow field;

[0019] The optical flow constraint equation is solved to obtain an optical flow field result for dynamic area detection.

[0020] Preferably, the detection of the dynamic area in step 2 adopts a threshold method, wherein the determination conditions of the dynamic area are as follows:

[0021]

[0022] Where M(x, y) represents the segmentation mask of the dynamic area, δ is the threshold, and u and v represent the horizontal and vertical displacement of the pixel in the optical flow field;

[0023] If M(x, y) = 1, it means that the pixel belongs to the dynamic area;

[0024] If M(x, y)=0, it means that the pixel belongs to the static background area.

[0025] Preferably, in step 3, the dynamic background recovery uses the Poisson equation for depth compensation, and the Poisson equation is defined as: KD(x, y) = 0,

[0026] Among them, K is the Laplace operator, which is used to calculate the second-order derivative of the pixel gradient.

[0027] D(x, y) represents the depth value that needs to be restored in the dynamic area.

[0028] Preferably, when the dynamic background recovery in step 3 uses the Poisson equation for depth compensation, the depth value of the recovery area is constrained by the following boundary conditions:

[0029] D(x, y) = D b (x, y),

[0030] (x,y)∈OR d ,

[0031] Among them, D(x, y) represents the depth value that needs to be restored in the dynamic area, D b (x, y) represents the known depth value of the dynamic region boundary, OR d Indicates the boundaries of the dynamic region.

[0032] Preferably, the edge weight calculation formula in step 4 is as follows:

[0033] w ij =exp(-B(∥I i -I j ∥ 2 +∥x i -x j ∥ 2 )),

[0034] Among them, w ij is the weight between pixels, I i ,I j is the pixel color information, x i is the pixel space position, and B is the parameter that controls the degree of similarity attenuation.

[0035] Preferably, the pixel feature space graph model constructed based on edge weights in step 4 further uses a conditional random field for label optimization, and the energy function of the conditional random field is defined as follows:

[0036] E(X)=∑ i∈V D i (x i )+∑ (i,j)∈E w ij ·S(x i , x j ),

[0037] Among them, X represents the set of labels to be optimized, V represents the set of pixels and voxel nodes, E represents the set of edges, and D i (x i ) is the label x i The matching cost between pixel i, S(x i , x j ) is the smoothing cost between labels, w ij Represents the edge weight between pixels and voxels.

[0038] Preferably, the label optimization of the conditional random field in step 4 is solved by a graph cut algorithm, and the optimal solution of the label optimization is expressed by the following formula:

[0039]

[0040] Among them, X * represents the optimized optimal label set, E(X) is the conditional random field energy function;

[0041] The graph cut algorithm refines the boundaries between the dynamic area and the static area by iteratively updating the label set, and provides the optimization result to the dense three-dimensional model construction in step 5.

[0042] Preferably, the dense three-dimensional model in step 5 is constructed using an octree structure for spatial division, and the relationship between the node depth and spatial resolution of the octree division is defined by the following formula:

[0043]

[0044] Among them, r is the spatial resolution of the current octree node, R is the spatial resolution of the root node, and d is the depth of the current division of the octree.

[0045] Preferably, the dense 3D model construction in step 5 is further combined with an incremental map update method to perform real-time optimization on the dynamic area and the static background, and the cost function of the incremental update is defined as follows:

[0046]

[0047] in, and Indicates the node information in the new and old maps, w ij represents the edge weight between nodes, Y is the smoothing factor, E up (M) represents the cost function of incremental map update, M represents the node set of the 3D model,

[0048] Represents the square of the Euclidean distance between the three-dimensional information of adjacent nodes i and j after update.

[0049] The present invention provides a method for real-time spatial reconstruction based on deep learning. It has the following beneficial effects:

[0050] 1. The present invention realizes accurate segmentation and positioning of dynamic objects through a dynamic object detection method based on optical flow constraints, combining the temporal changes of pixel points and the motion characteristics of dynamic areas with depth data analysis, thereby significantly improving the accuracy of dynamic object detection in dynamic environments.

[0051] 2. The present invention uses the Poisson equation to constrain the boundary information of the dynamic area and guide the depth compensation, thereby realizing the geometric continuity restoration of the depth value inside the dynamic area, and achieving the effect of improving the consistency of dynamic background compensation and the accuracy of depth information.

[0052] 3. The present invention optimizes the global consistency of dynamic and static regions through a label propagation mechanism based on conditional random fields and graph models, thereby improving the continuity of dynamic and static region boundaries and the accuracy of label assignment, and achieving the effect of enhancing the spatiotemporal consistency of label segmentation results in dynamic scenarios.

[0053] 4. The present invention adopts an incremental map update method combined with edge weight optimization and node depth compensation to achieve dynamic adjustment and update of the map frame by frame, thereby maintaining the geometric accuracy and significantly improving the stability and real-time performance of map updates in dynamic scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION

[0055] In order to make the technical personnel in the technical field understand the scheme of the present invention, the technical scheme in the embodiment of the present invention will be clearly and completely described below in combination with the drawings in the embodiment of the present invention. Obviously, the described embodiment is a partial embodiment of the present invention, not a complete embodiment. Based on the embodiment of the present invention, other embodiments obtained by ordinary technicians in the field without creative work should fall within the scope of protection of the present invention.

[0056] The present invention is described in detail below in conjunction with the accompanying drawings:

[0057] Example:

[0058] Please see attached Figure 1 , an embodiment of the present invention provides a spatial real-time reconstruction method based on deep learning, comprising:

[0059] Step 1: prepare input data, including obtaining input image sequence and corresponding depth data and camera pose information, obtaining camera intrinsic parameter matrix for transforming image coordinates to camera coordinates, and performing distortion correction and denoising on input image sequence and depth data, and providing the processed image sequence and depth data to step 2 for dynamic object detection and analysis;

[0060] Step 2, dynamic object detection, including calculating the time change information and spatial gradient information of the pixel points based on the image sequence and depth data processed in step 1, extracting the optical flow field by analyzing the displacement change of the pixel points between time frames, analyzing the motion characteristics of the dynamic object area in combination with the depth data, generating a dynamic segmentation mask by setting a threshold to distinguish the dynamic area from the static background area, and passing the generated dynamic segmentation mask and the position information of the dynamic area as input to step 3 for dynamic background recovery;

[0061] Step 3, dynamic background recovery, including determining the dynamic area and the boundary area based on the dynamic segmentation mask generated in step 2 and the depth data provided in step 1, analyzing the boundary pixel information of the dynamic area in the dynamic segmentation mask, constraining the depth compensation of the dynamic area according to the depth value of the boundary area, guiding the recovery of the internal depth value of the dynamic area in combination with the continuity of the boundary depth information, generating a depth map after dynamic background recovery, and providing the restored depth map and the dynamic segmentation mask to step 4 for global consistency modeling and optimization;

[0062] Step 4: Global consistency modeling, including building a graph model of the pixel feature space based on the depth map after dynamic background restoration in step 3 and the dynamic segmentation mask generated in step 2, characterizing the feature differences between dynamic and static regions through edge weights, optimizing the global consistency of dynamic and static regions through a label propagation mechanism, updating pixel and voxel labels, and using the optimized labels and pixel features to describe the spatial distribution of the static background and dynamic foreground of the scene. The optimization results and the depth map after background restoration are provided to step 5 for building a dense 3D model.

[0063] Step 5: dense 3D model construction and update, including generating a dense 3D model based on the pixel features and label information optimized in step 4 and the depth map after dynamic background restoration in step 3, dynamically adjusting the model content frame by frame through an incremental map update method, performing depth compensation on the occluded area in combination with the background restoration result in step 3, and using a spatial division structure to manage the storage and spatial partition calculation of the dense model, and finally outputting the dense 3D model as a complete reconstruction result of the static background, and performing unified coordinate correction on the dense model according to the posture information provided in step 1.

[0064] Benefits of step 1: By acquiring image sequences, depth data, and camera pose information, distortion correction and denoising are performed to ensure the quality and consistency of input data, providing an input data source for subsequent dynamic object detection and analysis. At the same time, the camera intrinsic parameter matrix is ​​used for coordinate transformation to ensure accurate mapping of spatial information in different coordinate systems, laying the foundation for dynamic object detection and depth compensation;

[0065] Benefits of step 2: Based on the optical flow field, the temporal changes and spatial gradient information of the pixels are calculated to effectively identify the dynamic object area, and its motion characteristics are analyzed in combination with the depth data to further generate a dynamic segmentation mask. The dynamic segmentation mask separates the dynamic object from the static background, reduces the interference of the dynamic object on the spatial modeling, provides clear boundaries and area divisions for dynamic background recovery, and ensures the accuracy and continuity of subsequent background recovery;

[0066] Benefits of step 3: The Poisson equation is used to constrain the depth information of the dynamic area boundary, guide the recovery of the internal depth value, realize the depth compensation and background recovery of the dynamic area, ensure the geometric continuity and depth consistency of the dynamic occluded area, effectively overcome the reconstruction incompleteness problem caused by background loss and occlusion in dynamic scenes, and provide depth map input for subsequent global consistency modeling;

[0067] Benefits of step 4: Use the pixel feature space graph model to analyze the feature differences between dynamic and static areas, optimize label distribution through label propagation mechanism, enhance the consistency of labels in time and space, and ensure the smoothness and continuity of segmentation boundaries. At the same time, solve the problems of label drift and boundary discontinuity in dynamic environments, and provide globally consistent and optimized segmentation results and label information for the construction of dense 3D models;

[0068] Benefits of step 5: The model content is adjusted frame by frame through the incremental map update method, and depth compensation is performed in combination with the dynamic background recovery results to ensure real-time update and geometric accuracy of the map model in a dynamic environment. In addition, the spatial division structure is used to optimize the storage and computing burden, improve the processing efficiency and real-time performance of the model in large-scale scenes, and finally output the complete static background reconstruction result after unified coordinate correction to ensure the accuracy and consistency of the model.

[0069] The calculation of the optical flow field in step 2 is based on the following optical flow constraint equation:

[0070] I x u+I y v+I t =0,

[0071] Among them, I x ,I y is the spatial gradient of the image in the horizontal and vertical directions, I t is the temporal gradient of the image, u and v represent the horizontal and vertical displacement of pixels in the optical flow field;

[0072] The optical flow constraint equation is solved to obtain the optical flow field result, which is used for dynamic area detection.

[0073] The present invention adopts an optical flow field calculation method in dynamic object detection, extracts the motion characteristics and displacement information of pixels through optical flow constraint equations, realizes high-precision segmentation of dynamic areas, and effectively solves the problems of misjudgment of dynamic object detection in complex environments, discontinuous segmentation boundaries, and insufficient real-time performance. At the same time, dynamic area input information is provided for subsequent steps to ensure the continuity and accuracy of background recovery and global consistency modeling. The optical flow field calculation method of the present invention has the advantages of high accuracy, strong adaptability, good real-time performance, etc., and provides a basis for real-time reconstruction of space in dynamic environments.

[0074] The detection of dynamic areas in step 2 adopts a threshold method, where the determination conditions of dynamic areas are as follows:

[0075]

[0076] Where M(x, y) represents the segmentation mask of the dynamic area, δ is the threshold, and u and v represent the horizontal and vertical displacement of the pixel in the optical flow field;

[0077] If M(x, y) = 1, it means that the pixel belongs to the dynamic area;

[0078] If M(x, y)=0, it means that the pixel belongs to the static background area.

[0079] The present invention adopts a threshold-based determination method in dynamic area detection, and realizes the segmentation of dynamic area and static background by analyzing the horizontal and vertical displacement of optical flow field pixels and comparing the displacement amplitude with the set threshold. The method is simple and efficient in calculation, has high detection accuracy and environmental adaptability, provides dynamic area segmentation results for subsequent dynamic background restoration and global consistency modeling, and ensures real-time requirements.

[0080] The threshold detection method reduces the misjudgment rate of dynamic area detection in complex environments, ensures the connection and consistency of dynamic masks in subsequent processing steps, and lays a data foundation for real-time spatial reconstruction in dynamic environments.

[0081] In step 3, the dynamic background recovery uses the Poisson equation for depth compensation, which is defined as:

[0082] KD(x,y)=0,

[0083] Among them, K is the Laplace operator, which is used to calculate the second-order derivative of the pixel gradient.

[0084] D(x, y) represents the depth value that needs to be restored in the dynamic area.

[0085] The present invention uses the Poisson equation for depth compensation in dynamic background restoration, calculates the gradient change of the dynamic area through the Laplace operator, and uses the boundary information to constrain the internal depth distribution of the dynamic area to ensure that the restored depth value has continuity and consistency in space. This method can accurately compensate for the missing depth information of the dynamic occlusion area, overcome the problems of uneven boundary transition and geometric distortion in traditional interpolation methods, and provide a background restoration solution for complex dynamic scenes.

[0086] The application of Poisson's equation in dynamic background recovery enhances the geometric accuracy of depth compensation, ensures the fusion effect of dynamic background information and static background, provides input data for subsequent global consistency optimization and construction of dense three-dimensional models, and significantly improves the accuracy and completeness of real-time spatial reconstruction in dynamic environments.

[0087] When the dynamic background recovery in step 3 uses the Poisson equation for depth compensation, the depth value of the recovery area is constrained by the following boundary conditions:

[0088] D(x, y) = D b (x, y),

[0089] (x,y)∈OR d ,

[0090] Among them, D(x, y) represents the depth value that needs to be restored in the dynamic area, D b (x, y) represents the known depth value of the dynamic region boundary, OR d Indicates the boundaries of the dynamic region.

[0091] The present invention combines the Poisson equation and boundary conditions to constrain the depth value of the restored area in dynamic background restoration, guides the depth compensation process of the restored area through the known depth value of the boundary, and ensures the geometric continuity and boundary consistency between the restored result and the boundary information.

[0092] This method effectively solves the problem of missing depth information due to occlusion in dynamic environments, maintains smooth boundary transition and internal structure continuity during the restoration process, and provides a more robust solution for the accurate restoration of occluded areas of dynamic objects in complex scenes. At the same time, this method provides depth data input for subsequent global consistency optimization and dense 3D model construction, significantly improving the overall accuracy and stability of real-time spatial reconstruction in dynamic environments.

[0093] The edge weight calculation formula in step 4 is as follows:

[0094] w ij =exp(-B(∥I i -I j ∥ 2 +∥x i -x j ∥2 )),

[0095] Among them, w ij is the weight between pixels, I i ,I j is the pixel color information, x i is the pixel space position, and B is the parameter that controls the degree of similarity attenuation.

[0096] The present invention adopts a method of calculating edge weights based on pixel color information and spatial position in global consistency modeling, controls the degree of similarity attenuation by edge weights, effectively characterizes the feature differences between dynamic areas and static backgrounds, and improves boundary detection accuracy and label propagation consistency.

[0097] This method utilizes the sensitivity of edge weights to color and spatial changes to adapt to scene changes in complex dynamic environments, ensures the smoothness of label assignment and boundary optimization in the global modeling process, provides input conditions for real-time spatial reconstruction in dynamic scenes, and enhances the robustness and adaptability of the model, thereby improving the accuracy and stability of the final 3D reconstruction results.

[0098] The pixel feature space graph model constructed based on edge weights in step 4 is further optimized using conditional random fields. The energy function of the conditional random field is defined as follows:

[0099] E(X)=∑ i∈V D i (x i )+∑ (i,j)∈E w ij ·S(x i , x j ),

[0100] Among them, X represents the set of labels to be optimized, V represents the set of pixels and voxel nodes, E represents the set of edges, and D i (x i ) is the label x i The matching cost between pixel i, S(x i , x j ) is the smoothing cost between labels, w ij Represents the edge weight between pixels and voxels.

[0101] The present invention constructs a pixel feature space graph model based on edge weights in global consistency modeling, further adopts conditional random fields to optimize label allocation, and improves the accuracy of label optimization and the smoothness of boundaries by defining matching cost and smoothing cost to minimize the energy function.

[0102] This method effectively solves the problems of unstable label division and boundary discontinuity in dynamic environments, enhances the robustness of label optimization through the sensitivity of edge weights to feature differences, ensures that the segmentation results of dynamic areas and static backgrounds in complex dynamic scenes have high spatial consistency and temporal stability, and provides accurate label input for the subsequent construction of dense three-dimensional models, thereby improving the overall performance and practicality of spatial real-time reconstruction.

[0103] The label optimization of the conditional random field in step 4 is solved based on the graph cut algorithm. The optimal solution of label optimization is expressed by the following formula:

[0104]

[0105] Among them, X * represents the optimized optimal label set, E(X) is the conditional random field energy function;

[0106] The graph cut algorithm iteratively updates the label set to refine the boundary between the dynamic area and the static area, and provides the optimization result to the dense 3D model construction in step 5.

[0107] The present invention adopts a solution method based on a graph cut algorithm in the label optimization process of a conditional random field, iteratively updates the label set through a global optimization mechanism, realizes the refinement of the boundary between the dynamic area and the static background, and ensures that the label division result has higher accuracy and continuity.

[0108] This method overcomes the problem of traditional label partitioning algorithms that are imprecise in processing boundaries and prone to errors, and further enhances the robustness and stability of label partitioning in dynamic scenes by combining the energy constraints of conditional random fields and the efficient solving ability of graph cut algorithms. At the same time, the optimization results provide input conditions for the subsequent construction of dense 3D models, significantly improve the geometric accuracy and reconstruction consistency of real-time spatial reconstruction, and provide technical support for 3D modeling in dynamic environments.

[0109] In step 5, the dense 3D model is constructed using an octree structure for spatial division. The relationship between the node depth and spatial resolution of the octree division is defined by the following formula:

[0110]

[0111] Among them, r is the spatial resolution of the current octree node, R is the spatial resolution of the root node, and d is the depth of the current division of the octree.

[0112] In step 5, the dense 3D model construction is further combined with the incremental map update method to optimize the dynamic area and static background in real time. The cost function of the incremental update is defined as follows:

[0113]

[0114] in, and Indicates the node information in the new and old maps, w ij represents the edge weight between nodes, Y is the smoothing factor, E up (M) represents the cost function of incremental map update, M represents the node set of the 3D model,

[0115] Represents the square of the Euclidean distance between the three-dimensional information of adjacent nodes i and j after update.

[0116] The present invention adopts an octree structure for space division in the process of building a dense three-dimensional model. Through the formula definition of node depth and spatial resolution, dynamic resolution control and spatial optimization management are realized, which effectively reduces storage and calculation complexity and enhances adaptability in large-scale scenarios.

[0117] At the same time, combined with the incremental map update method, by constructing a cost function based on edge weight and Euclidean distance, the model content is updated dynamically in real time, maintaining the geometric consistency and boundary smoothness of the dynamic area and static background, and effectively solving the problems of update delay and geometric accuracy loss in dynamic scenes.

[0118] The spatial division and incremental update method proposed in the present invention ensures the high precision and high integrity of the dense three-dimensional model, while improving the stability and efficiency of real-time processing of the model in a dynamic environment, providing technical support for the real-time spatial reconstruction of complex dynamic scenes.

[0119] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A spatial real-time reconstruction method based on deep learning, characterized in that: include: Step 1: prepare input data, including obtaining input image sequence and corresponding depth data and camera pose information, obtaining camera intrinsic parameter matrix for transforming image coordinates to camera coordinates, and performing distortion correction and denoising on input image sequence and depth data, and providing the processed image sequence and depth data to step 2 for dynamic object detection and analysis; Step 2, dynamic object detection, including calculating the time change information and spatial gradient information of the pixel points based on the image sequence and depth data processed in step 1, extracting the optical flow field by analyzing the displacement change of the pixel points between time frames, analyzing the motion characteristics of the dynamic object area in combination with the depth data, generating a dynamic segmentation mask by setting a threshold to distinguish the dynamic area from the static background area, and passing the generated dynamic segmentation mask and the position information of the dynamic area as input to step 3 for dynamic background recovery; Step 3, dynamic background recovery, including determining the dynamic area and the boundary area based on the dynamic segmentation mask generated in step 2 and the depth data provided in step 1, analyzing the boundary pixel information of the dynamic area in the dynamic segmentation mask, constraining the depth compensation of the dynamic area according to the depth value of the boundary area, guiding the recovery of the internal depth value of the dynamic area in combination with the continuity of the boundary depth information, generating a depth map after dynamic background recovery, and providing the restored depth map and the dynamic segmentation mask to step 4 for global consistency modeling and optimization; Step 4: Global consistency modeling, including building a graph model of the pixel feature space based on the depth map after dynamic background restoration in step 3 and the dynamic segmentation mask generated in step 2, characterizing the feature differences between dynamic and static regions through edge weights, optimizing the global consistency of dynamic and static regions through a label propagation mechanism, updating pixel and voxel labels, and using the optimized labels and pixel features to describe the spatial distribution of the static background and dynamic foreground of the scene. The optimization results and the depth map after background restoration are provided to step 5 for building a dense 3D model. Step 5: dense 3D model construction and update, including generating a dense 3D model based on the pixel features and label information optimized in step 4 and the depth map after dynamic background restoration in step 3, dynamically adjusting the model content frame by frame through an incremental map update method, performing depth compensation on the occluded area in combination with the background restoration result in step 3, and using a spatial division structure to manage the storage and spatial partition calculation of the dense model, and finally outputting the dense 3D model as a complete reconstruction result of the static background, and performing unified coordinate correction on the dense model according to the posture information provided in step 1.

2. The method for real-time spatial reconstruction based on deep learning according to claim 1, characterized in that: The calculation of the optical flow field in step 2 is based on the following optical flow constraint equation: I x u+I y v+I t =0, Among them, I x ,I y is the spatial gradient of the image in the horizontal and vertical directions, I t is the temporal gradient of the image, u and v represent the horizontal and vertical displacement of pixels in the optical flow field; The optical flow constraint equation is solved to obtain an optical flow field result for dynamic area detection.

3. The method for real-time spatial reconstruction based on deep learning according to claim 2, characterized in that: The detection of the dynamic area in step 2 adopts a threshold method, wherein the determination conditions of the dynamic area are as follows: Where M(x, y) represents the segmentation mask of the dynamic area, δ is the threshold, and u and v represent the horizontal and vertical displacement of the pixel in the optical flow field; If M(x, y) = 1, it means that the pixel belongs to the dynamic area; If M(x, y)=0, it means that the pixel belongs to the static background area.

4. The method for real-time spatial reconstruction based on deep learning according to claim 1, characterized in that: In step 3, the dynamic background recovery uses the Poisson equation for depth compensation, and the Poisson equation is defined as: KD(x, y) = 0, Among them, K is the Laplace operator, which is used to calculate the second-order derivative of the pixel gradient. D(x, y) represents the depth value that needs to be restored in the dynamic area.

5. The method for real-time spatial reconstruction based on deep learning according to claim 4, characterized in that: When the dynamic background recovery in step 3 uses the Poisson equation for depth compensation, the depth value of the recovery area is constrained by the following boundary conditions: D(x,y)=D b (x,y), (x, y)∈OR d , Among them, D(x, y) represents the depth value that needs to be restored in the dynamic area, D b (x, y) represents the known depth value of the dynamic region boundary, OR d Indicates the boundaries of the dynamic region.

6. The method for real-time spatial reconstruction based on deep learning according to claim 1, characterized in that: The edge weight calculation formula in step 4 is as follows: w ij =exp(-B(∥I i -I j ∥ 2 +∥x i -x j ∥ 2 )), Among them, w ij is the weight between pixels, I i ,I j is the pixel color information, x i is the pixel space position, and B is the parameter that controls the degree of similarity attenuation.

7. The method for real-time spatial reconstruction based on deep learning according to claim 6, characterized in that: The pixel feature space graph model constructed based on edge weights in step 4 further uses conditional random fields for label optimization. The energy function of the conditional random field is defined as follows: E(X)=∑ i∈V D i (x i )+∑ (i,j)∈E w ij ·S(x i ,x j ), Among them, X represents the set of labels to be optimized, V represents the set of pixels and voxel nodes, E represents the set of edges, and D i (x i ) is the label x i The matching cost between pixel i, S(x i , x j ) is the smoothing cost between labels, w ij Represents the edge weight between pixels and voxels.

8. The method for real-time spatial reconstruction based on deep learning according to claim 7, characterized in that: The label optimization of the conditional random field in step 4 is solved by a graph cut algorithm, and the optimal solution of the label optimization is expressed by the following formula: Among them, X * represents the optimized optimal label set, E(X) is the conditional random field energy function; The graph cut algorithm refines the boundaries between the dynamic area and the static area by iteratively updating the label set, and provides the optimization result to the dense three-dimensional model construction in step 5.

9. The method for real-time spatial reconstruction based on deep learning according to claim 1, characterized in that: In step 5, the dense 3D model is constructed using an octree structure for spatial division. The relationship between the node depth and spatial resolution of the octree division is defined by the following formula: Among them, r is the spatial resolution of the current octree node, R is the spatial resolution of the root node, and d is the depth of the current division of the octree.

10. The method for real-time spatial reconstruction based on deep learning according to claim 9, characterized in that: The dense 3D model construction in step 5 is further combined with the incremental map update method to perform real-time optimization on the dynamic area and static background. The cost function of the incremental update is defined as follows: in, and Indicates the node information in the new and old maps, w ij represents the edge weight between nodes, Y is the smoothing factor, E up (M) represents the cost function of incremental map update, M represents the node set of the 3D model, Represents the square of the Euclidean distance between the three-dimensional information of adjacent nodes i and j after update.