Three-dimensional reconstruction method, system and device based on monocular video and medium
By capturing video with a monocular camera and combining it with sliding window segmentation and local reconstruction technology, the problems of 3D reconstruction efficiency and accuracy in small space environments are solved, and high-precision real-time 3D reconstruction is achieved.
Patent Information
- Application Number
- CN202510722081.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-05-30
AI Technical Summary
Existing 3D reconstruction methods in small spatial environments suffer from registration ambiguity due to the increase in feature duplication areas caused by the narrow field of view, and the limited movement range reduces the effectiveness of motion parallax, affecting the efficiency and accuracy of 3D reconstruction.
Video is collected by a monocular camera, and sliding window segmentation processing and local reconstruction technology are used. Combined with key frame joint registration and spatial constraint optimization, the window length is adaptively adjusted to perform local reconstruction and global optimization to improve feature extraction and registration accuracy.
High-precision, high-integrity real-time 3D reconstruction is achieved in a small space environment, reducing monocular scale drift and registration errors, and improving the accuracy and efficiency of 3D reconstruction.
Smart Images

Figure CN120672942A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and in particular to a three-dimensional reconstruction method, system, device and medium based on monocular video. Background Art
[0002] In related technologies, 3D reconstruction methods typically use multiple stereo cameras and other devices to collect large amounts of 3D point clouds or image data, and then convert this data into a 3D model using a 3D reconstruction algorithm. However, in practical applications, it has been found that in small spatial environments, the narrow field of view of these methods increases the number of repeated feature areas, causing registration ambiguity. Furthermore, the limited range of motion reduces the effectiveness of motion parallax, exacerbating monocular scale drift and affecting the efficiency of 3D reconstruction. In summary, the technical problems existing in related technologies need to be improved. Summary of the Invention
[0003] The main purpose of the embodiments of the present application is to propose a three-dimensional reconstruction method, system, device and medium based on monocular video, which can improve the accuracy of three-dimensional reconstruction.
[0004] To achieve the above objectives, an embodiment of the present application provides a method for 3D reconstruction based on monocular video, the method comprising:
[0005] The monocular camera is used to collect and process data of the indoor space to obtain a monocular video;
[0006] Performing sliding window segmentation processing on the monocular video according to the spatial volume of the indoor space to obtain video clips;
[0007] Performing local reconstruction processing on the video clip to obtain a local reconstructed point cloud;
[0008] Performing key frame joint registration processing on the local reconstructed point cloud to obtain a registered scene frame;
[0009] A global scene optimization process is performed on the registered scene frame according to the spatial constraints to obtain a three-dimensional reconstruction result.
[0010] In some embodiments, performing sliding window segmentation processing on the monocular video according to the spatial volume of the indoor space to obtain video segments includes the following steps:
[0011] Performing volume calculation processing on the indoor space according to the monocular video to obtain the volume of the space;
[0012] Initializing a sliding window according to the spatial volume, and adjusting the length of the sliding window based on motion blur detection to obtain a target sliding window;
[0013] The monocular video is segmented according to the target sliding window to obtain the video clips.
[0014] In some embodiments, performing local reconstruction processing on the video clip to obtain a local reconstructed point cloud includes the following steps:
[0015] Performing feature extraction processing on the video clip through an image encoder, and performing temporal feature fusion processing on the extracted features using a gated recurrent unit to obtain spatial scene features;
[0016] Performing multi-view information fusion processing on the spatial scene features through a key frame decoder to obtain multi-view information;
[0017] Performing key frame information supplementation processing on the spatial scene features by supporting a frame decoder to obtain key frame information;
[0018] Performing bidirectional cross-attention calculation processing on the multi-view information and the key frame information according to spatial locality constraints to obtain fusion features;
[0019] The point cloud regression module based on deformable convolution performs regression prediction processing on the fused features to obtain the local reconstructed point cloud.
[0020] In some embodiments, performing keyframe joint registration processing on the local reconstructed point cloud to obtain a registered scene frame includes the following steps:
[0021] Acquire a scene frame buffer pool, wherein the scene frame buffer pool includes historical scene frames;
[0022] Performing coordinate transformation processing on the local reconstructed point cloud to obtain global point cloud data;
[0023] The global point cloud data is subjected to registration retrieval processing according to the scene frame buffer pool to obtain the registered scene frame.
[0024] In some embodiments, performing registration retrieval processing on the scene frame buffer pool according to the global point cloud data to obtain the registered scene frame includes the following steps:
[0025] Performing cosine similarity retrieval processing on each of the historical scene frames in the scene frame buffer pool according to the global point cloud data to generate a key frame set;
[0026] Performing spatiotemporal feature alignment processing on the key frame set to obtain cross-key frame spatiotemporal features;
[0027] Performing three-dimensional point cloud registration processing on the key frame set according to the cross-key frame spatiotemporal features to obtain a registered point cloud;
[0028] Performing point cloud fusion processing on the registered point cloud to obtain the registered scene frame.
[0029] In some embodiments, performing global scene optimization processing on the registered scene frame according to the spatial constraint to obtain a three-dimensional reconstruction result includes the following steps:
[0030] Performing point cloud optimization processing on the registered scene frame to obtain point cloud optimization data;
[0031] Performing plane constraint optimization processing on the point cloud optimization data to obtain plane optimization data;
[0032] Performing spatial topology optimization processing on the plane optimization data to obtain the three-dimensional reconstruction result.
[0033] In some embodiments, performing spatial topology optimization processing on the plane optimization data to obtain the three-dimensional reconstruction result includes the following steps:
[0034] Performing topology construction processing on the plane optimization data to obtain a scene topology graph;
[0035] The scene topology graph is iteratively optimized according to a graph convolutional network to obtain the three-dimensional reconstruction result.
[0036] To achieve the above objectives, another aspect of the present application provides a 3D reconstruction system based on monocular video, the system comprising:
[0037] The first module is used to acquire and process the video of the indoor space through a monocular camera to obtain a monocular video;
[0038] The second module is configured to perform sliding window segmentation processing on the monocular video according to the spatial volume of the indoor space to obtain video clips;
[0039] The third module is used to perform local reconstruction processing on the video clip to obtain a local reconstructed point cloud;
[0040] The fourth module is used to perform key frame joint registration processing on the local reconstructed point cloud to obtain a registered scene frame;
[0041] The fifth module is used to perform global scene optimization processing on the registered scene frame according to the spatial constraints to obtain a three-dimensional reconstruction result.
[0042] To achieve the above-mentioned purpose, another aspect of an embodiment of the present application provides an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned method when executing the computer program.
[0043] To achieve the above objectives, another aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described above is implemented.
[0044] The embodiments of the present application include at least the following beneficial effects: The present application provides a method, system, device, and medium for 3D reconstruction based on monocular video. This solution uses a monocular camera to collect and process data from an indoor space to obtain a monocular video, then performs sliding window segmentation processing on the monocular video based on the spatial volume of the indoor space to obtain video segments. The solution can adaptively adjust the window length based on the spatial volume, increase the local window overlap rate, reduce monocular scale drift, and provide a data foundation for subsequent local reconstruction. Furthermore, this solution performs local reconstruction processing on the video segments to obtain a local reconstructed point cloud, performs keyframe joint registration processing on the local reconstructed point cloud to obtain a registered scene frame, and performs global scene optimization processing on the registered scene frame based on spatial constraints to obtain a 3D reconstruction result. This solution can detect reconstruction completeness based on spatial constraints, reduce registration errors, and improve the accuracy of 3D reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 This is a flowchart of a 3D reconstruction method based on monocular video provided in an embodiment of the present application;
[0046] Figure 2 Schematic diagram of a three-dimensional reconstruction system based on monocular video provided in an embodiment of the present application;
[0047] Figure 3 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0048] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are merely examples of systems and methods consistent with some aspects of the embodiments of the present application as detailed in the appended claims.
[0049] It will be understood that the terms "first", "second", etc. used in this application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0050] The terms "at least one", "plurality", "each", "any", etc. used in this application include "at least one", "two" or more, "plurality" or "each", "any" or "any one", "each" or "any one" as used herein.
[0051] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0052] In the related art, the three-dimensional reconstruction method is usually based on a large amount of three-dimensional point cloud or image data collected by multiple stereo cameras and other equipment, and then the data is converted into a three-dimensional model through a three-dimensional reconstruction algorithm. However, in actual applications, it is found that the related method has an increase in feature duplication areas due to the narrow field of view in a small space environment, which leads to registration ambiguity, and the limited range of motion reduces the effectiveness of motion parallax, aggravates the monocular scale drift, and affects the efficiency of three-dimensional reconstruction. In summary, the technical problems existing in the related art need to be improved. For example, for example, the related art adopts the simultaneous localization and mapping (SLAM) method for three-dimensional reconstruction, but this method requires offline processing and cannot meet the real-time requirements, and the real-time dense SLAM system has deficiencies in reconstruction accuracy and integrity. The solution based on depth sensor is expensive and restricted by the environment.
[0053] In view of this, embodiments of the present application provide a method, system, device, and medium for 3D reconstruction based on monocular video. This solution uses a monocular camera to collect and process data from an indoor space to obtain a monocular video. This solution then performs sliding window segmentation processing on the monocular video based on the spatial volume of the indoor space to obtain video segments. This solution can adaptively adjust the window length based on the spatial volume, increase the local window overlap rate, reduce monocular scale drift, and provide a data foundation for subsequent local reconstruction. Furthermore, this solution performs local reconstruction processing on the video segments to obtain a local reconstructed point cloud, performs keyframe joint registration processing on the local reconstructed point cloud to obtain a registered scene frame, and performs global scene optimization processing on the registered scene frame based on spatial constraints to obtain a 3D reconstruction result. This solution can detect reconstruction completeness based on spatial constraints, reduce registration errors, and improve the accuracy of 3D reconstruction.
[0054] The embodiment of the present application provides a three-dimensional reconstruction method based on monocular video, which relates to the field of computer vision technology. The embodiment of the present application provides a three-dimensional reconstruction method based on monocular video, which can be applied to a terminal, a server, or a software running in a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, and a car terminal, etc., but is not limited to this; the server side can be configured as an independent physical server, or as a server cluster or distributed system composed of multiple physical servers, or as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application that implements a three-dimensional reconstruction method based on monocular video, etc., but is not limited to the above forms.
[0055] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0056] Figure 1 This is an optional flowchart of a 3D reconstruction method based on monocular video provided in an embodiment of the present application. Figure 1 The method may include but is not limited to steps S101 to S105.
[0057] Step S101, collecting and processing data of the indoor space through a monocular camera to obtain a monocular video;
[0058] Step S102, performing sliding window segmentation processing on the monocular video according to the spatial volume of the indoor space to obtain video segments;
[0059] Step S103, performing local reconstruction processing on the video clip to obtain a local reconstructed point cloud;
[0060] Step S104, performing key frame joint registration processing on the local reconstructed point cloud to obtain a registered scene frame;
[0061] Step S105 , performing global scene optimization processing on the registered scene frame according to the spatial constraint to obtain a three-dimensional reconstruction result.
[0062] In steps S101 to S105, an embodiment of the present application uses only an ordinary monocular camera to perform high-precision, high-integrity real-time 3D reconstruction in a small space environment. In this embodiment, the "small space" refers to an indoor enclosed space with an area of 10-50 square meters and a height of no more than 5 meters. Specifically, the embodiment of the present application uses a monocular camera to collect data from the indoor space to obtain a monocular video, then adaptively adjusts the length of the sliding window according to the spatial volume, and uses the sliding window to segment the monocular video, dividing the input video stream into overlapping short segments. The video segment is then locally reconstructed, and a local reconstructed point cloud is obtained using an improved multi-branch neural network model. The dense 3D point cloud image of each frame in the window can be directly predicted, and the intermediate frame is used as a key frame to establish a local coordinate system. The embodiment of the present application also performs key frame joint registration processing on the local reconstructed point cloud, incrementally aligning the local reconstructed point cloud to the global coordinate system. Based on visual similarity and baseline suitability scores, historical scene frames related to the current frame can be retrieved from the buffer pool, thereby obtaining a registered scene frame through joint registration. Finally, a global scene optimization process is performed on the registered scene frames based on the spatial constraints to obtain the 3D reconstruction result. This embodiment of the application specifically optimizes the following parameters to address the characteristics of small space environments: the window length is set to 11 frames to balance reconstruction quality and efficiency; the scene frame buffer pool size is set to 30-50 frames; and each registration retrieves the 5-10 most relevant scene frames, using a multi-keyframe co-registration strategy for registration.
[0063] One of the above technical solutions has the following advantages or beneficial effects: the embodiment of the present application can reduce registration ambiguity and monocular scale drift through dynamic window segmentation strategy and spatial topology constraints, thereby improving the accuracy of three-dimensional reconstruction.
[0064] In some embodiments, performing sliding window segmentation processing on the monocular video according to the spatial volume of the indoor space to obtain video segments includes the following steps:
[0065] Performing volume calculation processing on the indoor space according to the monocular video to obtain the volume of the space;
[0066] Initializing a sliding window according to the spatial volume, and adjusting the length of the sliding window based on motion blur detection to obtain a target sliding window;
[0067] The monocular video is segmented according to the target sliding window to obtain the video clips.
[0068] In the embodiment of the present application, the volume of the indoor space can be calculated by the monocular video obtained by acquisition, or the spatial volume can be obtained by actually measuring the indoor space. The embodiment of the present application can perform spatial recognition processing on the monocular video through a deep learning model, detect the wall or ground of the space in the monocular video, and predict the depth of the space through a depth prediction model, thereby calculating the volume based on the area and depth obtained by the detection to obtain the spatial volume. Then, a sliding window is initialized according to the spatial volume, for example, the window frame number V represents the spatial volume. The length of the sliding window is then adjusted based on motion blur detection. The degree of blur can be detected using image sharpness or edge information, and the window length is adaptively adjusted based on the blur level. The monocular video is then segmented using the adaptively adjusted target sliding window to produce video segments.
[0069] One of the above technical solutions has the following advantages or beneficial effects: The embodiment of the present application can adaptively adjust the window size according to scene requirements by dynamically adjusting the sliding window, thereby achieving an optimal balance between resources and performance, and providing a data basis for subsequent three-dimensional reconstruction.
[0070] In some embodiments, performing local reconstruction processing on the video clip to obtain a local reconstructed point cloud includes the following steps:
[0071] Performing feature extraction processing on the video clip through an image encoder, and performing temporal feature fusion processing on the extracted features using a gated recurrent unit to obtain spatial scene features;
[0072] Performing multi-view information fusion processing on the spatial scene features through a key frame decoder to obtain multi-view information;
[0073] Performing key frame information supplementation processing on the spatial scene features by supporting a frame decoder to obtain key frame information;
[0074] Performing bidirectional cross-attention calculation processing on the multi-view information and the key frame information according to spatial locality constraints to obtain fusion features;
[0075] The point cloud regression module based on deformable convolution performs regression prediction processing on the fused features to obtain the local reconstructed point cloud.
[0076] In the embodiment of the present application, the video clip is subjected to feature extraction processing by an image encoder. In order to introduce a cross-window feature transfer mechanism to address the high repetitiveness of features in small-space scenes, a gated recurrent unit is used to perform temporal feature fusion processing on the extracted features. The formula of the gated recurrent unit is as follows:
[0077] z t =σ(W z ·[h t-1 ,x t ])
[0078] Where z t represents the update gate output vector of the current time step, σ represents the Sigmoid activation function, W z represents the learnable weight matrix corresponding to the update gate, h t-1 represents the hidden state vector of the previous time step, x t Represents the input feature vector of the current time step. The key frame decoder then performs multi-view information fusion processing on the spatial scene features to obtain multi-view information, and the supporting frame decoder performs key frame information supplementation processing on the spatial scene features to obtain key frame information. The embodiment of the present application also designs a spatial locality constraint for the bidirectional cross attention mechanism, and limits the attention radius according to the characteristics of the small space environment. The calculation formula of the attention radius is as follows:
[0079]
[0080] Among them, r is the attention radius, W and H are the image width and height respectively. The embodiment of the present application forces attention to the local geometric associations in the small space scene through the attention radius. It should be noted that the embodiment of the present application uses the intermediate frame as the key frame to establish the local coordinate system. Finally, the point cloud regression module based on deformable convolution performs regression prediction processing on the fusion features to obtain a local reconstructed point cloud. Among them, when the embodiment of the present application adds a refinement module based on deformable convolution to the point cloud regression module, the convolution kernel deformation offset is limited according to the characteristics of the small space. The calculation formula of the convolution kernel deformation offset is as follows:
[0081]
[0082] Among them, Δp is the convolution kernel deformation offset, D is the scene depth estimation value, and f is the focal length parameter. The embodiment of the present application ensures that the local optimization in a small space meets the spatial scale constraint by limiting the convolution kernel deformation offset.
[0083] One of the above technical solutions has the following advantages or beneficial effects: The embodiment of the present application can better detect the features of a narrow field of view by locally reconstructing the video clips and adding corresponding spatial scale constraints in combination with the characteristics of the small space environment, thereby improving the accuracy of feature extraction and reconstruction.
[0084] In some embodiments, performing keyframe joint registration processing on the local reconstructed point cloud to obtain a registered scene frame includes the following steps:
[0085] Acquire a scene frame buffer pool, wherein the scene frame buffer pool includes historical scene frames;
[0086] Performing coordinate transformation processing on the local reconstructed point cloud to obtain global point cloud data;
[0087] The global point cloud data is subjected to registration retrieval processing according to the scene frame buffer pool to obtain the registered scene frame.
[0088] In an embodiment of the present application, the scene frame buffer pool is a pre-built buffer pool that includes multiple frames of historical scene frames. The historical scene frames can be obtained from a database or by pre-collecting and processing the space. In this embodiment of the present application, global point cloud data is obtained by incrementally registering the local reconstructed point cloud to the global coordinate system. The global point cloud data is then registered and retrieved based on the scene frame buffer pool. Feature reuse across windows is achieved through the pre-built scene frame buffer library. Cosine similarity is used to retrieve multiple similar historical scene frames as key frames for registration, thereby obtaining a registered scene frame.
[0089] One of the above technical solutions has the following advantages or beneficial effects: the embodiment of the present application can perform batch multi-keyframe joint registration by performing registration and retrieval processing on global point cloud data, thereby improving the efficiency of registration.
[0090] In some embodiments, performing registration retrieval processing on the scene frame buffer pool according to the global point cloud data to obtain the registered scene frame includes the following steps:
[0091] Performing cosine similarity retrieval processing on each of the historical scene frames in the scene frame buffer pool according to the global point cloud data to generate a key frame set;
[0092] Performing spatiotemporal feature alignment processing on the key frame set to obtain cross-key frame spatiotemporal features;
[0093] Performing three-dimensional point cloud registration processing on the key frame set according to the cross-key frame spatiotemporal features to obtain a registered point cloud;
[0094] Performing point cloud fusion processing on the registered point cloud to obtain the registered scene frame.
[0095] In an embodiment of the present application, the similarity of each historical scene frame in the scene frame buffer pool is calculated based on the cosine similarity of the global point cloud data. A similarity threshold can be set for screening to generate a key frame set. The key frame set is then aligned with the spatiotemporal features. The spatiotemporal feature cube can be constructed and the spatiotemporal feature cube can be extracted with a three-dimensional convolution kernel to obtain cross-keyframe spatiotemporal features. The key frame set is then processed for three-dimensional point cloud registration based on the cross-keyframe spatiotemporal features. The improved three-dimensional point cloud registration (ICP) algorithm is used for multi-keyframe joint registration. The objective function is:
[0096]
[0097] The weight w k Dynamically calculated by key frame confidence, R represents the rotation transformation matrix, represents the coordinates of the ith 3D point of the kth key frame in the source point cloud, t represents the translation transformation vector, Indicates the target point cloud The corresponding three-dimensional point coordinates. Finally, the registered point cloud is fused to obtain the registered scene frame. Specifically, point cloud fusion can be performed by establishing a probabilistic fusion model. The expression of the fusion model is as follows:
[0098]
[0099] In the formula, p(x) represents the fusion probability density function of the three-dimensional space point x, α krepresents the mixing weight coefficient of the kth Gaussian component, Represents the three-dimensional Gaussian distribution probability density function, μ k represents the mean vector of the kth Gaussian component, Σ k The embodiment of the present application can also solve the optimal fusion parameters by using the expectation maximization algorithm, and set the point cloud confidence threshold to 3 to filter low-quality reconstructed point cloud data.
[0100] One of the above technical solutions has the following advantages or beneficial effects: the embodiment of the present application adopts a multi-keyframe co-registration strategy, which can simultaneously perform registration processing on multiple keyframes, thereby improving the efficiency and accuracy of registration.
[0101] In some embodiments, performing global scene optimization processing on the registered scene frame according to the spatial constraint to obtain a three-dimensional reconstruction result includes the following steps:
[0102] Performing point cloud optimization processing on the registered scene frame to obtain point cloud optimization data;
[0103] Performing plane constraint optimization processing on the point cloud optimization data to obtain plane optimization data;
[0104] Performing spatial topology optimization processing on the plane optimization data to obtain the three-dimensional reconstruction result.
[0105] In the embodiments of the present application, a small space optimization strategy is also set up based on the characteristics of small space scenes. By introducing point cloud distribution optimization, the registered scene frame is subjected to point cloud optimization processing to obtain point cloud optimization data. Plane constraint optimization is then performed on the point cloud optimization data to obtain plane optimization data. For example, by adding plane constraint optimization to large planes such as walls for constraint optimization, spatial topology optimization can also be introduced to perform spatial topology optimization on the plane optimization data to obtain a three-dimensional reconstruction result. Specifically, point cloud distribution optimization uses a density-aware clustering algorithm to optimize the point cloud data by defining a density metric, where the expression of the density metric is as follows:
[0106]
[0107] Where ρ(x) represents the density metric, x represents the coordinates of the target three-dimensional point to be calculated, and x i Represents the coordinates of the i-th adjacent point in the neighborhood N(x). Then, by setting the adaptive neighborhood radius r = μ d +ασ d , where μ d is the average neighbor distance, α represents, σ dRepresented. The distribution of point cloud data is optimized based on density metric and adaptive neighborhood radius. The embodiment of the present application also uses a multi-plane detection algorithm for plane constraint optimization, wherein the expression of the plane detection algorithm is as follows:
[0108]
[0109] Where n represents the plane normal vector, d represents the distance from the plane to the origin, The embodiment of the present application establishes a plane relationship graph and enforces orthogonal constraints, such as an orthogonal constraint of 90±5° between wall surfaces, thereby constraining the point cloud data to obtain optimized plane optimization data.
[0110] One of the above technical solutions has the following advantages or beneficial effects: the embodiment of the present application can improve the accuracy of three-dimensional reconstruction by introducing plane orthogonality constraint optimization and density-adaptive topology optimization.
[0111] In some embodiments, performing spatial topology optimization processing on the plane optimization data to obtain the three-dimensional reconstruction result includes the following steps:
[0112] Performing topology construction processing on the plane optimization data to obtain a scene topology graph;
[0113] The scene topology graph is iteratively optimized according to a graph convolutional network to obtain the three-dimensional reconstruction result.
[0114] In the embodiment of the present application, a scene topology graph G = (V, E, W) is obtained by constructing plane optimization data, where vertices represent spatial units. Then, the point cloud distribution is optimized by a graph convolutional network (GCN). The optimization formula is as follows:
[0115]
[0116] Where H (l+1) represents the node feature matrix of the (l+1)th layer, σ represents a nonlinear activation function, and the embodiment of the present application adopts the ReLU activation function, represents the normalized form of the degree matrix, represents the normalized form of the adjacency matrix, H (l) represents the input node feature of the lth layer, W (l) Represents the trainable weight matrix of layer l.
[0117] One of the above technical solutions has the following advantages or beneficial effects: The embodiment of the present application performs spatial topology optimization on point cloud data through a graph convolutional network, which can make the reconstructed data more suitable for a small space environment and improve the accuracy of three-dimensional reconstruction.
[0118] The following is a detailed description of the embodiments of the present application with reference to specific application examples:
[0119] The embodiment of the present application receives a monocular RGB video input, initializes the first window, tries all frames as key frame candidates, selects the reconstruction result with the highest total confidence to initialize the global scene, and optimizes by establishing a small space prior constraint to obtain a three-dimensional reconstruction result. The embodiment of the present application obtains video clips through a sliding window, extracts features of each frame through an image encoder, fuses multi-view information through a key frame decoder, supplements key frame information through a support frame decoder, and predicts 3D point clouds and confidence through a regression head. The reconstructed point cloud data is then aligned to the global coordinate system, the relevant historical scene frames are retrieved through the scene frame buffer pool, the coordinate system of the decoder is transformed by jointly encoding the image and geometric features, the global scene is optimized through the scene decoder, and the scene frame buffer pool is updated at the same time. The optimization process is then performed through a small space optimization strategy, and a plane orthogonality constraint is introduced to eliminate wall registration errors. Specifically, the embodiment of the present application adopts a two-level neural network framework, divides the video into short clips through a sliding window mechanism, directly predicts local 3D point clouds using the first-level network, and then incrementally aligns them to the global coordinate system through the second-level network. Specifically for small environments, the system optimizes window size, scene frame management strategies, and spatial constraints, achieving high-quality real-time reconstruction without explicit camera parameter estimation. By introducing volume-aware window segmentation, orthogonality-constrained optimization, and density-adaptive topology optimization, the system achieves a 42.7% improvement in reconstruction completeness and a reduction in registration error to 0.11m in a standard 5m×5m test scene compared to other 3D reconstruction methods, all while maintaining real-time performance of 23 FPS.
[0120] See also Figure 2 The present application also provides a monocular video-based 3D reconstruction system, which can implement the above-mentioned monocular video-based 3D reconstruction method. The system includes:
[0121] The first module 201 is used to obtain a monocular video of the indoor space through a monocular camera;
[0122] The second module 202 is configured to perform sliding window segmentation processing on the monocular video according to the spatial volume of the indoor space to obtain video segments;
[0123] The third module 203 is configured to perform local reconstruction processing on the video clip to obtain a local reconstructed point cloud;
[0124] The fourth module 204 is configured to perform key frame joint registration processing on the local reconstructed point cloud to obtain a registered scene frame;
[0125] The fifth module 205 is configured to perform global scene optimization processing on the registered scene frame according to the spatial constraints to obtain a three-dimensional reconstruction result.
[0126] It can be understood that the contents of the above method embodiments are all applicable to the present system embodiments, the functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0127] The present application also provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described monocular video-based 3D reconstruction method. The electronic device can be any smart terminal, including a tablet computer and an in-vehicle computer.
[0128] It can be understood that the contents of the above method embodiments are applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0129] See also Figure 3 , Figure 3 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes:
[0130] The processor 301 may be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.
[0131] The memory 302 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 302 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 302 and is called by the processor 301 to execute the monocular video-based 3D reconstruction method of the embodiments of this application.
[0132] Input / output interface 303, used to implement information input and output;
[0133] Communication interface 304, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);
[0134] bus 305 , which transmits information between the various components of the device (e.g., processor 301 , memory 302 , input / output interface 303 , and communication interface 304 );
[0135] The processor 301 , the memory 302 , the input / output interface 303 and the communication interface 304 are connected to each other in communication within the device via the bus 305 .
[0136] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned three-dimensional reconstruction method based on monocular video.
[0137] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiment, the functions specifically implemented by the present storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0138] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0139] The embodiments of the present application provide a monocular video-based 3D reconstruction method, system, device, and medium. This solution uses a monocular camera to collect and process data from an indoor space to obtain a monocular video. This solution then performs sliding window segmentation processing on the monocular video based on the spatial volume of the indoor space to obtain video segments. This method can adaptively adjust the window length based on the spatial volume, increase the local window overlap rate, reduce monocular scale drift, and provide a data foundation for subsequent local reconstruction. Furthermore, this solution performs local reconstruction processing on the video segments to obtain a local reconstructed point cloud, performs keyframe joint registration processing on the local reconstructed point cloud to obtain a registered scene frame, and performs global scene optimization processing on the registered scene frame based on spatial constraints to obtain a 3D reconstruction result. This solution can detect reconstruction completeness based on spatial constraints, reduce registration errors, and improve the accuracy of 3D reconstruction.
[0140] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0141] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0142] The system embodiment described above is merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0143] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.
[0144] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0145] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0146] In the several embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the above units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of systems or units, which can be electrical, mechanical or other forms.
[0147] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0148] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0149] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0150] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.
Claims
1. A three-dimensional reconstruction method based on monocular video, characterized in that: The method comprises the following steps: The monocular camera is used to collect and process data of the indoor space to obtain a monocular video; Performing sliding window segmentation processing on the monocular video according to the spatial volume of the indoor space to obtain video clips; Performing local reconstruction processing on the video clip to obtain a local reconstructed point cloud; Performing key frame joint registration processing on the local reconstructed point cloud to obtain a registered scene frame; A global scene optimization process is performed on the registered scene frame according to the spatial constraints to obtain a three-dimensional reconstruction result.
2. The method according to claim 1, characterized in that The step of performing sliding window segmentation processing on the monocular video according to the spatial volume of the indoor space to obtain video segments comprises the following steps: Performing volume calculation processing on the indoor space according to the monocular video to obtain the volume of the space; Initializing a sliding window according to the spatial volume, and adjusting the length of the sliding window based on motion blur detection to obtain a target sliding window; The monocular video is segmented according to the target sliding window to obtain the video clips.
3. The method according to claim 1, characterized in that The locally reconstructing the video clip to obtain a locally reconstructed point cloud comprises the following steps: Performing feature extraction processing on the video clip through an image encoder, and performing temporal feature fusion processing on the extracted features using a gated recurrent unit to obtain spatial scene features; Performing multi-view information fusion processing on the spatial scene features through a key frame decoder to obtain multi-view information; Performing key frame information supplementation processing on the spatial scene features by supporting a frame decoder to obtain key frame information; Performing bidirectional cross-attention calculation processing on the multi-view information and the key frame information according to spatial locality constraints to obtain fusion features; The point cloud regression module based on deformable convolution performs regression prediction processing on the fused features to obtain the local reconstructed point cloud.
4. The method according to claim 1, wherein The step of performing key frame joint registration processing on the local reconstructed point cloud to obtain a registered scene frame includes the following steps: Acquire a scene frame buffer pool, wherein the scene frame buffer pool includes historical scene frames; Performing coordinate transformation processing on the local reconstructed point cloud to obtain global point cloud data; The global point cloud data is subjected to registration retrieval processing according to the scene frame buffer pool to obtain the registered scene frame.
5. The method according to claim 4, characterized in that The step of performing registration retrieval processing on the scene frame buffer pool according to the global point cloud data to obtain the registered scene frame comprises the following steps: Performing cosine similarity retrieval processing on each of the historical scene frames in the scene frame buffer pool according to the global point cloud data to generate a key frame set; Performing spatiotemporal feature alignment processing on the key frame set to obtain cross-key frame spatiotemporal features; Performing three-dimensional point cloud registration processing on the key frame set according to the cross-key frame spatiotemporal features to obtain a registered point cloud; Performing point cloud fusion processing on the registered point cloud to obtain the registered scene frame.
6. The method according to claim 1, characterized in that The step of performing global scene optimization processing on the registered scene frame according to the spatial constraints to obtain a three-dimensional reconstruction result comprises the following steps: Performing point cloud optimization processing on the registered scene frame to obtain point cloud optimization data; Performing plane constraint optimization processing on the point cloud optimization data to obtain plane optimization data; Performing spatial topology optimization processing on the plane optimization data to obtain the three-dimensional reconstruction result.
7. The method according to claim 6, characterized in that The performing of spatial topology optimization processing on the plane optimization data to obtain the three-dimensional reconstruction result comprises the following steps: Performing topology construction processing on the plane optimization data to obtain a scene topology graph; The scene topology graph is iteratively optimized according to a graph convolutional network to obtain the three-dimensional reconstruction result.
8. A three-dimensional reconstruction system based on monocular video, characterized in that: The system comprises: The first module is used to acquire and process the video of the indoor space through a monocular camera to obtain a monocular video; The second module is configured to perform sliding window segmentation processing on the monocular video according to the spatial volume of the indoor space to obtain video clips; The third module is used to perform local reconstruction processing on the video clip to obtain a local reconstructed point cloud; The fourth module is used to perform key frame joint registration processing on the local reconstructed point cloud to obtain a registered scene frame; The fifth module is used to perform global scene optimization processing on the registered scene frame according to the spatial constraints to obtain a three-dimensional reconstruction result.
9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Image processing device and image processing method
CN104103062A
Accurate burn area calculation method based on three-dimensional human body reconstruction
CN110310285A
Unmanned aerial vehicle scene dense reconstruction method based on VI-SLAM and depth estimation network
CN112435325A
Indoor three-dimensional reconstruction and whole design method and system based on big data
CN118470203A
Three-dimensional mapping method and system
CN119206106A