A three-dimensional scene reconstruction method and system based on monocular depth estimation
Through a three-dimensional scene reconstruction method based on monocular depth estimation, using artificial intelligence algorithms to optimize feature extraction and depth estimation, the problems of high cost, insufficient accuracy and low automation in large-scale boundless indoor environments are solved, and efficient and stable three-dimensional reconstruction is achieved.
Patent Information
- Application Number
- CN202510600126.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-05-12
AI Technical Summary
The prior art has problems such as high cost, insufficient accuracy, low automation and unstable reconstruction results in three-dimensional reconstruction of large-scale boundless indoor environments.
A three-dimensional scene reconstruction method based on monocular depth estimation is adopted, and a reconstruction scheme generation model, a two-dimensional image processing model and a three-dimensional scene reconstruction model are constructed using artificial intelligence algorithms. Two-dimensional images are collected through a monocular camera and depth estimation and three-dimensional reconstruction are carried out. Feature extraction and depth estimation are optimized in combination with LSTM-ISGA, CNN-MOGRPO-IPA and FPN-ResNet-3D-DBN algorithms.
It reduces equipment costs, improves the accuracy and automation of monocular depth estimation, enhances the stability and adaptability of reconstruction results, and adapts to three-dimensional reconstruction under complex textures and variable lighting conditions.
Smart Images

Figure CN120147554B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of three-dimensional scene reconstruction, and in particular relates to a three-dimensional scene reconstruction method and system based on monocular depth estimation. Background Art
[0002] Large-scale unbounded indoor environments generally refer to extremely large indoor spaces without distinct divisions or boundaries. These environments are typically characterized by vast expanses, lack of clear boundaries, and multifunctionality. With the widespread adoption of digital technology, the digital transformation of these large-scale unbounded indoor environments has become a key research topic. With the rapid development of technologies such as virtual reality, augmented reality, and intelligent robotics, 3D scene reconstruction technology has shown broad application prospects in the digital transformation of large-scale unbounded indoor environments.
[0003] However, existing technologies still face many challenges in 3D reconstruction and virtualization of large-scale unbounded indoor environments, including the following deficiencies:
[0004] 1) High cost: Existing technologies often rely on multi-camera or depth cameras for 3D reconstruction. These devices are expensive, limiting the popularization and application of the technology.
[0005] 2) Insufficient precision: The accuracy of 3D reconstruction using monocular depth estimation in existing technologies is relatively low, especially in areas with complex textures, large lighting changes, or no textures, where the accuracy of depth estimation is severely affected.
[0006] 3) Low degree of automation: The existing technologies often require manual intervention in the process of shooting point setting, image processing, and 3D reconstruction, resulting in a low degree of automation.
[0007] 4) Unstable reconstruction results: Due to the uncertainty of depth estimation, the 3D reconstruction results of existing technologies are often unstable and easily affected by noise and outliers. Summary of the Invention
[0008] In order to solve the problems of high cost investment, insufficient accuracy, low degree of automation and unstable reconstruction results in the existing technology, the purpose of the present invention is to provide a three-dimensional scene reconstruction method and system based on monocular depth estimation.
[0009] The technical solution adopted in the present invention is:
[0010] A three-dimensional scene reconstruction method based on monocular depth estimation comprises the following steps:
[0011] Use artificial intelligence algorithms to build reconstruction solution generation models, 2D image processing models, and 3D scene reconstruction models;
[0012] Based on the real-time indoor layout data of a large-scale unbounded indoor environment, a reconstruction solution generation model is used to generate a reconstruction solution and obtain a real-time reconstruction solution.
[0013] According to the real-time reconstruction scheme, a monocular camera is used to collect multiple real-time 2D images of a large-scale unbounded indoor environment from multiple perspectives.
[0014] According to the real-time reconstruction scheme, a two-dimensional image processing model is used to process a plurality of real-time two-dimensional images to obtain a plurality of processed real-time two-dimensional images;
[0015] According to the real-time reconstruction scheme, a three-dimensional scene reconstruction model is used to perform depth estimation and three-dimensional scene reconstruction on a number of processed real-time two-dimensional images to obtain real-time three-dimensional scene data.
[0016] Furthermore, the real-time reconstruction solution includes real-time two-dimensional image acquisition decision, real-time two-dimensional image processing effect requirement decision, and real-time three-dimensional scene data reconstruction effect requirement decision.
[0017] Furthermore, artificial intelligence algorithms are used to construct a reconstruction solution generation model, a two-dimensional image processing model, and a three-dimensional scene reconstruction model, including the following steps:
[0018] Collect a number of historical indoor layout data, and use a swarm intelligence optimization algorithm to build a reconstruction scheme generation model based on the historical indoor layout data to obtain a number of historical reconstruction schemes; the historical reconstruction schemes include historical 2D image acquisition decisions, historical 2D image processing effect requirements decisions, and historical 3D scene reconstruction effect requirements decisions;
[0019] Collecting a number of historical two-dimensional images, and constructing a two-dimensional image processing model using reinforcement learning and image processing algorithms based on the historical two-dimensional images and the historical two-dimensional image processing effects of a number of historical reconstruction schemes, to obtain a number of processed historical two-dimensional images;
[0020] Based on the historical three-dimensional scene reconstruction effect requirements of several processed historical two-dimensional images and several historical reconstruction schemes, a three-dimensional scene reconstruction model is constructed using deep learning and three-dimensional reconstruction algorithms.
[0021] Furthermore, the reconstruction scheme generation model is constructed based on the LSTM-ISGA algorithm, and the reconstruction scheme generation model includes an indoor layout feature extraction module constructed based on the LSTM algorithm and a reconstruction scheme generation module constructed based on the ISGA algorithm, which are connected in sequence. The reconstruction scheme generation module includes an initial solution generation module, an iterative optimization module and an optimal solution analysis module, which are connected in sequence.
[0022] Furthermore, the two-dimensional image processing model is constructed based on the CNN-MOGRPO-IPA algorithm, and the two-dimensional image processing model includes a two-dimensional image feature extraction module constructed based on the CNN algorithm, a processing strategy generation module constructed based on the MOGRPO algorithm, and a two-dimensional image processing algorithm library equipped with several IPA algorithms, which are connected in sequence.
[0023] Furthermore, the three-dimensional scene reconstruction model is constructed based on the FPN-ResNet-3D-DBN algorithm, and the three-dimensional scene reconstruction model includes a multi-scale fusion feature extraction module constructed based on the FPN algorithm, a depth estimation module constructed based on the ResNet algorithm, and a three-dimensional scene reconstruction module constructed based on the 3D-DBN algorithm, which are connected in sequence.
[0024] Furthermore, based on the real-time indoor layout data of the large-scale unbounded indoor environment, a reconstruction solution generation model is used to generate a reconstruction solution to obtain a real-time reconstruction solution, including the following steps:
[0025] Use the indoor layout feature extraction module of the reconstruction solution generation model to extract real-time indoor layout features of real-time indoor layout data of large-scale unbounded indoor environments and input them into the reconstruction solution generation module;
[0026] Using the initial solution generation module of the reconstruction solution generation module, an initial solution is generated according to the real-time indoor layout characteristics to obtain several initial solutions; the initial solutions correspond to the initial real-time reconstruction solutions;
[0027] Use the iterative optimization module of the reconstruction solution generation module to iteratively optimize several initial solutions and obtain the optimal solution with the best fitness value;
[0028] The optimal solution analysis module of the reconstruction solution generation module is used to analyze the individual vectors of the optimal solution to obtain the optimal real-time reconstruction solution.
[0029] Furthermore, according to the real-time reconstruction scheme, a two-dimensional image processing model is used to process a plurality of real-time two-dimensional images to obtain a plurality of processed real-time two-dimensional images, including the following steps:
[0030] Using a two-dimensional image feature extraction module of a two-dimensional image processing model to extract real-time two-dimensional image features of the real-time two-dimensional image;
[0031] Using a processing strategy generation module of a two-dimensional image processing model, a processing strategy is generated according to the real-time two-dimensional image processing effect requirement decision of a real-time reconstruction scheme and the real-time two-dimensional image characteristics to obtain a real-time two-dimensional image processing strategy;
[0032] According to the real-time two-dimensional image processing strategy, the corresponding target IPA algorithm is called in the two-dimensional image processing algorithm library, and the corresponding real-time two-dimensional image is processed according to the target IPA algorithm to obtain the corresponding processed real-time two-dimensional image;
[0033] Traverse all real-time two-dimensional images and repeat the above steps to obtain a processed real-time two-dimensional image corresponding to each real-time two-dimensional image.
[0034] Furthermore, according to the real-time reconstruction scheme, a 3D scene reconstruction model is used to perform depth estimation and 3D scene reconstruction on a plurality of processed real-time 2D images to obtain real-time 3D scene data, including the following steps:
[0035] Adjusting the model parameters of the 3D scene reconstruction model according to the real-time 3D scene data reconstruction effect requirement decision of the real-time reconstruction solution to obtain an adjusted 3D scene reconstruction model;
[0036] extracting real-time multi-scale fusion features of each processed real-time two-dimensional image using the adjusted multi-scale fusion feature extraction module of the adjusted three-dimensional scene reconstruction model;
[0037] Using the adjusted depth estimation module of the adjusted 3D scene reconstruction model, perform depth estimation based on each real-time multi-scale fusion feature to obtain a corresponding real-time depth map;
[0038] An adjusted 3D scene reconstruction module of the adjusted 3D scene reconstruction model is used to reconstruct the 3D scene according to all the real-time depth maps to obtain real-time 3D scene data.
[0039] A three-dimensional scene reconstruction system based on monocular depth estimation is used to implement a three-dimensional scene reconstruction method. The system includes a model construction unit, a reconstruction scheme generation unit, a two-dimensional image processing unit, and a depth estimation and three-dimensional scene reconstruction unit connected in sequence.
[0040] The beneficial effects of the present invention are:
[0041] This paper proposes a 3D scene reconstruction method and system based on monocular depth estimation. By using a monocular camera instead of a multi-camera or depth camera, the equipment cost investment is greatly reduced, making the 3D reconstruction technology more economical and affordable. The 2D feature extraction and depth estimation processes are optimized by an artificial intelligence algorithm, which effectively improves the accuracy of monocular depth estimation, especially the performance in complex textured and textureless areas. The reconstruction scheme generation model, the 2D image processing model and the 3D scene reconstruction model are constructed to realize automatic reconstruction scheme generation, 2D image processing and 3D scene reconstruction, and realize the full automation from shooting point setting to 3D reconstruction, reducing manual intervention, improving efficiency, adapting to the reconstruction scheme of large-scale unbounded indoor environment changes, and enhancing the versatility and flexibility of the technology. The 2D image processing model generated by the adaptive processing strategy and the 3D scene reconstruction model optimized by multi-scale fusion feature extraction and depth estimation can maintain high robustness and accuracy under variable lighting and complex texture conditions, thereby improving the stability and reliability of the 3D reconstruction results.
[0042] Other beneficial effects of the present invention will be further described in the specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 This is a flowchart of the three-dimensional scene reconstruction method based on monocular depth estimation in the present invention.
[0044] Figure 2 It is a structural block diagram of the three-dimensional scene reconstruction system based on monocular depth estimation in the present invention. DETAILED DESCRIPTION
[0045] The present invention will be further explained below with reference to the accompanying drawings and specific embodiments.
[0046] Example 1:
[0047] like Figure 1 As shown, this embodiment provides a three-dimensional scene reconstruction method based on monocular depth estimation, comprising the following steps:
[0048] S1: Using artificial intelligence algorithms, we build a reconstruction solution generation model, a 2D image processing model, and a 3D scene reconstruction model, including the following steps:
[0049] S1-1: Collect some historical indoor layout data and use a swarm intelligence optimization algorithm to build a reconstruction solution generation model based on the data to obtain several historical reconstruction solutions. The historical reconstruction solutions include historical 2D image acquisition decisions, historical 2D image processing effect requirements, and historical 3D scene reconstruction effect requirements.
[0050] The reconstruction solution generation model is based on the Long Short-Term Memory (LSTM)-Improved Snow Geese Algorithm (ISGA) algorithm. It consists of a sequentially connected indoor layout feature extraction module based on the LSTM algorithm and a reconstruction solution generation module based on the ISGA algorithm. The reconstruction solution generation module includes a sequentially connected initial solution generation module, an iterative optimization module, and an optimal solution analysis module.
[0051] LSTM can capture the spatiotemporal dependencies in indoor layout data and effectively extract key features of indoor layouts, including size, shape, and layout position. This effectively preserves the spatiotemporal information in the indoor layout data and provides accurate basic features for subsequent reconstruction scheme generation. The ISGA algorithm has a strong global search capability and can find near-optimal reconstruction schemes. It can automatically adjust the reconstruction scheme according to different indoor layouts and has strong adaptability.
[0052] S1-2: Collect several historical 2D images and, based on the historical 2D images and the historical 2D image processing effect requirements of several historical reconstruction schemes, use reinforcement learning and image processing algorithms to build a 2D image processing model to obtain several processed historical 2D images;
[0053] The two-dimensional image processing model is built based on the Convolutional Neural Network (CNN)-Multi-Objective Group Relative Policy Optimization (MOGRPO)-Image Processing Algorithm (IPA) algorithm. The two-dimensional image processing model includes a two-dimensional image feature extraction module built based on the CNN algorithm, a processing strategy generation module built based on the MOGRPO algorithm, and a two-dimensional image processing algorithm library equipped with several IPA algorithms.
[0054] The 2D image feature extraction module captures low-level features such as edges, textures, and shapes in images through multi-layer convolution and pooling operations. It is highly efficient and can quickly and accurately identify key information in the image as the basis for generating processing strategies. The processing strategy generation module can simultaneously consider multiple optimization goals, such as processing speed, accuracy, and robustness, to generate a comprehensive optimal processing strategy. It can automatically adjust and optimize the processing strategy according to different image characteristics and requirements, and has strong adaptability. The 2D image processing algorithm library provides a variety of image processing algorithms, such as filtering, enhancement, segmentation, morphological processing, and perspective alignment processing, which can meet the processing requirements of different scenarios and needs. It can flexibly configure and select appropriate image processing algorithms to achieve personalized processing and provide high-quality image data for subsequent 3D scene reconstruction.
[0055] The processing strategy generation module includes an objective function set, a policy network, an experience replay pool, and an intelligent agent. The intelligent agent is connected to the objective function set and the policy network respectively. The intelligent agent learns historical processing strategies through the experience replay pool and continuously optimizes its own policy generation capabilities. The intelligent agent controls the policy network based on the learned experience to generate more effective processing strategies. The design of the experience replay pool and the intelligent agent enables the model to continuously learn and optimize, improving the quality of policy generation. The processing strategy generation module uses a group exploration method to avoid falling into local optimal solutions to a certain extent. The policy network outputs the distribution probability of actions under a given state. The processing strategy generation module directly updates the policy network through gradients, eliminating the critic model in traditional reinforcement learning and making the algorithm structure more concise.
[0056] S1-3: Based on the requirements for historical 3D scene reconstruction effects from several processed historical 2D images and several historical reconstruction schemes, a 3D scene reconstruction model is constructed using deep learning and 3D reconstruction algorithms.
[0057] The 3D scene reconstruction model is built based on the Feature Pyramid Network (FPN)-Residual Network (ResNet)-Three-Dimensional Deep Belief Network (3D-DBN) algorithm. The 3D scene reconstruction model includes a multi-scale fusion feature extraction module based on the FPN algorithm, a depth estimation module based on the ResNet algorithm, and a 3D scene reconstruction module based on the 3D-DBN algorithm.
[0058] The FPN of the multi-scale fusion feature extraction module combines the top-down path and the bottom-up path to achieve the fusion of features of different scales, capture image information from fine-grained to coarse-grained, effectively integrate image features of different scales, enhance the ability of feature representation, and facilitate the accuracy of subsequent depth estimation. It not only retains the detailed information of the image, but also integrates the global semantic information, providing a more comprehensive feature basis for depth estimation. The fusion of multi-scale features helps to improve the adaptability of the depth estimation module to complex scenes and reduce the estimation error caused by scale changes; the ResNet of the depth estimation module solves the problem by introducing the residual learning mechanism. It solves the gradient vanishing problem in deep neural network training, achieves more accurate depth estimation, helps speed up network training, reduces training time, and is more robust to complex scenes with different lighting and textures, ensuring the stability of depth estimation. The 3D scene reconstruction module is trained with a large number of historical depth maps, learning to predict depth maps to 3D scene data. 3D-DBN is good at processing 3D data and can generate high-quality 3D scene models, restore the three-dimensional structure of the scene, and better restore detailed information in the scene, such as the shape and position relationship of objects. It is suitable for 3D scene reconstruction in various complex large-scale unbounded indoor environments.
[0059] S2: Based on the real-time indoor layout data of a large-scale unbounded indoor environment, a reconstruction solution generation model is used to generate a reconstruction solution to obtain a real-time reconstruction solution, including the following steps:
[0060] S2-1: Use the indoor layout feature extraction module of the reconstruction solution generation model to extract real-time indoor layout features of the real-time indoor layout data of a large-scale unbounded indoor environment and input them into the reconstruction solution generation module;
[0061] S2-2: Using the initial solution generation module of the reconstruction solution generation module, an initial solution is generated based on the real-time indoor layout characteristics to obtain several initial solutions; the initial solutions correspond to the initial real-time reconstruction solutions;
[0062] The formula is:
[0063]
[0064] Where, The initial ISGA individual generated by the Circle chaotic map sequence, that is, the initial solution; is a randomly generated ISGA individual; iis the individual indicator of ISGA; compared with the randomly distributed population, the initial position distribution of the improved ISGA population is more uniform when the initial population is generated by the Circle chaotic mapping sequence, which expands the search range of the ISGA population in space and increases the diversity of group positions. To a certain extent, it improves the defect that the algorithm is prone to falling into local extreme values, thereby improving the optimization efficiency of the algorithm;
[0065] S2-3: Use the iterative optimization module of the reconstruction solution generation module to iteratively optimize several initial solutions to obtain the optimal solution with the best fitness value, including the following steps:
[0066] S2-3-1: Several initial solutions are used as the initial ISGA population; each initial ISGA individual in the initial ISGA population corresponds to an initial solution;
[0067] S2-3-2: Use the fitness function and the iterative update module to obtain the initial fitness value of each initial ISGA individual in the initial ISGA population, and use the initial ISGA individual with the lowest fitness value as the leader goose;
[0068] The formula is:
[0069]
[0070] Where, is the real-time fitness function; is the 2D image acquisition cost function; is a two-dimensional image processing function; Reconstruction cost for 3D scene data; For ISGA individuals; are the first weight value, the second weight value, and the third weight value;
[0071] S2-3-3: Entering the exploration phase, the leader goose rotation mechanism, the calling guidance mechanism, and the dynamic reverse mechanism are introduced to iteratively update the initial ISGA population to obtain an updated ISGA population and retain the optimal individual;
[0072] The leader goose rotation mechanism selects a new leader goose in each iteration based on the fitness value of the ISGA individuals. This mechanism can prevent the leader goose from falling into the local optimum too early and enhance the global search capability of the algorithm.
[0073] The formula is:
[0074]
[0075] Where, Be the leader of an update; For the The third-to-last initial ISGA individual with the lowest fitness value in the initial ISGA population at the number of iterations; For the The fifth initial ISGA individual with the lowest fitness value in the initial ISGA population of the iteration number; is the current iteration number; is the optimal individual; is the first weight factor; is a random number; To be the initial leader;
[0076] The calling guidance mechanism uses the sound wave propagation attenuation model to adjust the position update of the individual ISGA according to the distance between the individual and the leader goose. The position update of the ISGA individual that is closer is more affected by the leader goose, and it can quickly approach the optimal solution. The position update of the ISGA individual that is farther away is less affected by the leader goose, and it can maintain a certain exploration ability. This mechanism can avoid excessive aggregation or dispersion of the group and improve the local search accuracy of the algorithm.
[0077] The formula is:
[0078]
[0079] Where, is an updated ISGA individual; For the The number of iterations of the initial ISGA individual; is the initial ISGA individual Received sound intensity; For the The number of iterations of the initial ISGA individual Corresponding sound intensity parameters; is the initial sound intensity; is the minimum acceptable sound intensity; is the convergence factor; is the initial ISGA individual with the farthest distance; is a random parameter; is the Brownian motion function; is the Brownian motion parameter; is the XOR processing symbol;
[0080]
[0081] Where, is the convergence factor; tanh(.) is the hyperbolic tangent function; is the current iteration number; is the maximum number of iterations; a max 、 amin are the maximum and minimum values of the convergence factor respectively; λ is the deceleration rate parameter, is the decrement period parameter, λ =-2 π , = π ;
[0082] Dynamic reverse mechanism, which dynamically reverses the initial ISGA individuals to improve the diversity of exploration directions and avoid falling into local optimality;
[0083] The formula is:
[0084]
[0085] Where, is an updated reverse ISGA individual; γ is the decreasing inertia coefficient; L max 、 L min are the maximum and minimum values of the vector space respectively;
[0086] The leader goose of one update, several ISGA individuals of one update and several reverse ISGA individuals of one update are integrated to obtain an ISGA population of one update, and the ISGA individual with the lowest fitness value is retained as the optimal individual;
[0087] S2-3-4: Entering the development stage, introducing the abnormal boundary strategy and Gaussian mutation mechanism, performing a second update on the once-updated ISGA population, obtaining a second-updated ISGA population, and retaining the optimal individual;
[0088] Abnormal boundary strategy, calculates the difference between the fitness value of each updated ISGA individual and the average fitness value of the group. For ISGA individuals whose fitness value is much higher than the group average, their position update method will be adjusted, such as using Gaussian mutation mechanism, larger step size or smaller step size. This mechanism can help individuals avoid falling into local optimality and improve the convergence speed and accuracy of the algorithm;
[0089] The formula is:
[0090]
[0091] Where, is the second updated ISGA individual; is an updated ISGA individual; is the fitness function; is the average fitness value of the group; is the ISGA individual with the highest fitness value; are the second weight factor and the third weight factor; is the Gaussian mutation mechanism parameter;
[0092] S2-3-5: If the number of iterations is greater than or equal to the iteration threshold or the fitness value of the optimal individual is less than the fitness threshold, the optimal individual is output as the optimal solution;
[0093] S2-4: Use the optimal solution analysis module of the reconstruction solution generation module to analyze the individual vectors of the optimal solution to obtain the optimal real-time reconstruction solution;
[0094] The real-time reconstruction solution includes real-time 2D image acquisition decision, real-time 2D image processing effect requirement decision, and real-time 3D scene data reconstruction effect requirement decision;
[0095] Real-time 2D image acquisition decisions include the number of real-time monocular cameras, the number of real-time shooting points, the real-time monocular camera posture, the real-time monocular camera shooting time, etc.
[0096] The decision on the real-time 2D image processing effect requirements includes the real-time image quality requirements, real-time image viewing angle requirements, real-time image size requirements, etc. of the processed real-time 2D image;
[0097] The decision on the effect of real-time 3D scene data reconstruction includes the requirements for the real-time data type, real-time granularity, and real-time accuracy of the real-time 3D scene data.
[0098] S3: Based on the real-time reconstruction scheme, a monocular camera is used to collect multiple real-time 2D images of a large-scale unbounded indoor environment from multiple perspectives.
[0099] S4: According to the real-time reconstruction scheme, a two-dimensional image processing model is used to process the plurality of real-time two-dimensional images to obtain a plurality of processed real-time two-dimensional images, including the following steps:
[0100] S4-1: using a two-dimensional image feature extraction module of a two-dimensional image processing model to extract real-time two-dimensional image features of a real-time two-dimensional image;
[0101] S4-2: Using the processing strategy generation module of the two-dimensional image processing model, a processing strategy is generated based on the real-time two-dimensional image processing effect requirement decision of the real-time reconstruction solution and the real-time two-dimensional image characteristics to obtain a real-time two-dimensional image processing strategy;
[0102] S4-3: According to the real-time two-dimensional image processing strategy, a corresponding target IPA algorithm is called in the two-dimensional image processing algorithm library, and the corresponding real-time two-dimensional image is processed according to the target IPA algorithm to obtain a corresponding processed real-time two-dimensional image;
[0103] S4-4: traverse all real-time two-dimensional images and repeat the above steps to obtain a processed real-time two-dimensional image corresponding to each real-time two-dimensional image;
[0104] S5: According to the real-time reconstruction scheme, using the 3D scene reconstruction model, depth estimation and 3D scene reconstruction are performed on the processed real-time 2D images to obtain real-time 3D scene data, including the following steps:
[0105] S5-1: adjusting model parameters of the 3D scene reconstruction model according to the real-time 3D scene data reconstruction effect requirement decision of the real-time reconstruction solution to obtain an adjusted 3D scene reconstruction model;
[0106] S5-2: using the adjusted multi-scale fusion feature extraction module of the adjusted 3D scene reconstruction model, extracting real-time multi-scale fusion features of each processed real-time 2D image;
[0107] S5-3: Using the adjusted depth estimation module of the adjusted 3D scene reconstruction model, perform depth estimation based on each real-time multi-scale fusion feature to obtain a corresponding real-time depth map;
[0108] S5-4: Using the adjusted 3D scene reconstruction module of the adjusted 3D scene reconstruction model, perform 3D scene reconstruction according to all real-time depth maps to obtain real-time 3D scene data.
[0109] Example 2:
[0110] like Figure 2 As shown, this embodiment provides a three-dimensional scene reconstruction system based on monocular depth estimation, which is used to implement a three-dimensional scene reconstruction method. The system includes a model construction unit, a reconstruction scheme generation unit, a two-dimensional image processing unit, and a depth estimation and three-dimensional scene reconstruction unit connected in sequence;
[0111] A model building unit, used to use artificial intelligence algorithms to build a reconstruction solution generation model, a two-dimensional image processing model, and a three-dimensional scene reconstruction model;
[0112] A reconstruction solution generation unit is used to generate a reconstruction solution based on the real-time indoor layout data of the large-scale unbounded indoor environment using a reconstruction solution generation model to obtain a real-time reconstruction solution;
[0113] A two-dimensional image processing unit is used to process a plurality of real-time two-dimensional images acquired by the monocular camera using a two-dimensional image processing model according to a real-time reconstruction scheme to obtain a plurality of processed real-time two-dimensional images;
[0114] The depth estimation and three-dimensional scene reconstruction unit is used to perform depth estimation and three-dimensional scene reconstruction on a number of processed real-time two-dimensional images using a three-dimensional scene reconstruction model according to a real-time reconstruction scheme to obtain real-time three-dimensional scene data.
[0115] This paper proposes a 3D scene reconstruction method and system based on monocular depth estimation. By using a monocular camera instead of a multi-camera or depth camera, the equipment cost investment is greatly reduced, making the 3D reconstruction technology more economical and affordable. The 2D feature extraction and depth estimation processes are optimized by an artificial intelligence algorithm, which effectively improves the accuracy of monocular depth estimation, especially the performance in complex textured and textureless areas. The reconstruction scheme generation model, the 2D image processing model and the 3D scene reconstruction model are constructed to realize automatic reconstruction scheme generation, 2D image processing and 3D scene reconstruction, and realize the full automation from shooting point setting to 3D reconstruction, reducing manual intervention, improving efficiency, adapting to the reconstruction scheme of large-scale unbounded indoor environment changes, and enhancing the versatility and flexibility of the technology. The 2D image processing model generated by the adaptive processing strategy and the 3D scene reconstruction model optimized by multi-scale fusion feature extraction and depth estimation can maintain high robustness and accuracy under variable lighting and complex texture conditions, thereby improving the stability and reliability of the 3D reconstruction results.
[0116] The present invention is not limited to the above optional embodiments. Anyone can derive various other forms of products based on the teachings of the present invention. The above specific embodiments should not be construed as limiting the scope of protection of the present invention. The scope of protection of the present invention shall be based on the scope defined in the claims, and the description can be used to interpret the claims.
Claims
1. A three-dimensional scene reconstruction method based on monocular depth estimation, characterized by: The steps include: Use artificial intelligence algorithms to build reconstruction solution generation models, 2D image processing models, and 3D scene reconstruction models; The reconstruction solution generation model is constructed based on the LSTM-ISGA algorithm, and the reconstruction solution generation model includes an indoor layout feature extraction module constructed based on the LSTM algorithm and a reconstruction solution generation module constructed based on the ISGA algorithm, which are connected in sequence. The reconstruction solution generation module includes an initial solution generation module, an iterative optimization module, and an optimal solution analysis module, which are connected in sequence. The two-dimensional image processing model is constructed based on the CNN-MOGRPO-IPA algorithm, and the two-dimensional image processing model includes a two-dimensional image feature extraction module constructed based on the CNN algorithm, a processing strategy generation module constructed based on the MOGRPO algorithm, and a two-dimensional image processing algorithm library provided with several IPA algorithms, which are connected in sequence; The three-dimensional scene reconstruction model is constructed based on the FPN-ResNet-3D-DBN algorithm, and the three-dimensional scene reconstruction model includes a multi-scale fusion feature extraction module constructed based on the FPN algorithm, a depth estimation module constructed based on the ResNet algorithm, and a three-dimensional scene reconstruction module constructed based on the 3D-DBN algorithm, which are connected in sequence; Based on the real-time indoor layout data of a large-scale unbounded indoor environment, a reconstruction solution generation model is used to generate a reconstruction solution and obtain a real-time reconstruction solution. The real-time reconstruction solution includes real-time two-dimensional image acquisition decision, real-time two-dimensional image processing effect requirement decision and real-time three-dimensional scene data reconstruction effect requirement decision; According to the real-time reconstruction scheme, a monocular camera is used to collect multiple real-time 2D images of a large-scale unbounded indoor environment from multiple perspectives. According to the real-time reconstruction scheme, a two-dimensional image processing model is used to process a plurality of real-time two-dimensional images to obtain a plurality of processed real-time two-dimensional images; According to the real-time reconstruction scheme, a 3D scene reconstruction model is used to perform depth estimation and 3D scene reconstruction on a number of processed real-time 2D images to obtain real-time 3D scene data, including the following steps: Adjusting the model parameters of the 3D scene reconstruction model according to the real-time 3D scene data reconstruction effect requirement decision of the real-time reconstruction solution to obtain an adjusted 3D scene reconstruction model; extracting real-time multi-scale fusion features of each processed real-time two-dimensional image using the adjusted multi-scale fusion feature extraction module of the adjusted three-dimensional scene reconstruction model; Using the adjusted depth estimation module of the adjusted 3D scene reconstruction model, perform depth estimation based on each real-time multi-scale fusion feature to obtain a corresponding real-time depth map; An adjusted 3D scene reconstruction module of the adjusted 3D scene reconstruction model is used to reconstruct the 3D scene according to all the real-time depth maps to obtain real-time 3D scene data.
2. The method for 3D scene reconstruction based on monocular depth estimation according to claim 1, wherein: Using artificial intelligence algorithms, we build a reconstruction solution generation model, a 2D image processing model, and a 3D scene reconstruction model, including the following steps: Collecting a number of historical indoor layout data, and using a swarm intelligence optimization algorithm based on the number of historical indoor layout data, constructing a reconstruction scheme generation model to obtain a number of historical reconstruction schemes; the historical reconstruction schemes include historical two-dimensional image acquisition decisions, historical two-dimensional image processing effect requirements decisions, and historical three-dimensional scene reconstruction effect requirements decisions; Collecting a number of historical two-dimensional images, and constructing a two-dimensional image processing model using reinforcement learning and image processing algorithms based on the historical two-dimensional images and the historical two-dimensional image processing effects of a number of historical reconstruction schemes, to obtain a number of processed historical two-dimensional images; Based on the historical three-dimensional scene reconstruction effect requirements of several processed historical two-dimensional images and several historical reconstruction schemes, a three-dimensional scene reconstruction model is constructed using deep learning and three-dimensional reconstruction algorithms.
3. The method for 3D scene reconstruction based on monocular depth estimation according to claim 2, wherein: Based on the real-time indoor layout data of a large-scale unbounded indoor environment, a reconstruction solution generation model is used to generate a reconstruction solution to obtain a real-time reconstruction solution, including the following steps: Use the indoor layout feature extraction module of the reconstruction solution generation model to extract real-time indoor layout features of real-time indoor layout data of large-scale unbounded indoor environments and input them into the reconstruction solution generation module; Using the initial solution generation module of the reconstruction solution generation module, an initial solution is generated according to the real-time indoor layout characteristics to obtain a plurality of initial solutions; the initial solutions correspond to the initial real-time reconstruction solutions; Use the iterative optimization module of the reconstruction solution generation module to iteratively optimize several initial solutions and obtain the optimal solution with the best fitness value; The optimal solution analysis module of the reconstruction solution generation module is used to analyze the individual vectors of the optimal solution to obtain the optimal real-time reconstruction solution.
4. The method for 3D scene reconstruction based on monocular depth estimation according to claim 3, wherein: According to the real-time reconstruction scheme, a two-dimensional image processing model is used to process a plurality of real-time two-dimensional images to obtain a plurality of processed real-time two-dimensional images, including the following steps: Using a two-dimensional image feature extraction module of a two-dimensional image processing model to extract real-time two-dimensional image features of the real-time two-dimensional image; Using a processing strategy generation module of a two-dimensional image processing model, a processing strategy is generated according to the real-time two-dimensional image processing effect requirement decision of a real-time reconstruction scheme and the real-time two-dimensional image characteristics to obtain a real-time two-dimensional image processing strategy; According to the real-time two-dimensional image processing strategy, the corresponding target IPA algorithm is called in the two-dimensional image processing algorithm library, and the corresponding real-time two-dimensional image is processed according to the target IPA algorithm to obtain the corresponding processed real-time two-dimensional image; Traverse all real-time two-dimensional images and repeat the above steps to obtain a processed real-time two-dimensional image corresponding to each real-time two-dimensional image.
5. A 3D scene reconstruction system based on monocular depth estimation, for implementing the 3D scene reconstruction method according to any one of claims 1 to 4, characterized in that: The system comprises a model building unit, a reconstruction scheme generating unit, a two-dimensional image processing unit and a depth estimation and three-dimensional scene reconstruction unit which are connected in sequence.
Citation Information
Patent Citations
Three-dimensional reconstruction method based on attention mechanism and monocular multi-view angle
CN113838191A
Three-dimensional scene reconstruction method and system based on mixed reality glasses and application
CN114219900A