Three-dimensional scene reconstruction method and system based on monocular depth estimation

By adopting a three-dimensional scene reconstruction method based on monocular depth estimation in a large-scale bounded indoor environment, and using artificial intelligence algorithms to achieve automated processing, the problems of high cost, low accuracy, low degree of automation and unstable reconstruction results in the existing technology are solved, and efficient, economical and automated three-dimensional reconstruction effects are achieved.

CN120147554AActive Publication Date: 2025-06-13YUNHAI SPACETIME (BEIJING) TECH CO LTD +1

Patent Information

Application Number
CN202510600126.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-06-13
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

The prior art has problems such as high cost, low accuracy, low automation and unstable reconstruction results in three-dimensional reconstruction of large-scale boundless indoor environments.

Method used

A three-dimensional scene reconstruction method based on monocular depth estimation is adopted to construct reconstruction scheme generation model, two-dimensional image processing model and three-dimensional scene reconstruction model through artificial intelligence algorithms, so as to realize automated reconstruction scheme generation, two-dimensional image processing and three-dimensional scene reconstruction.

Benefits of technology

It reduces equipment cost investment, improves the accuracy of monocular depth estimation, realizes full automation from shooting point setting to three-dimensional reconstruction, enhances the universality and flexibility of technology, and improves the stability and reliability of the three-dimensional reconstruction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147554A_ABST
    Figure CN120147554A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of three-dimensional scene reconstruction, and discloses a three-dimensional scene reconstruction method and system based on monocular depth estimation. The method comprises the following steps: constructing a reconstruction scheme generation model, a two-dimensional image processing model and a three-dimensional scene reconstruction model; generating a reconstruction scheme by using a reconstruction scheme generation model according to the real-time indoor layout data to obtain a real-time reconstruction scheme; acquiring a plurality of real-time two-dimensional images by using a monocular camera according to the real-time reconstruction scheme; according to the real-time reconstruction scheme, processing the plurality of real-time two-dimensional images by using a two-dimensional image processing model to obtain a plurality of processed real-time two-dimensional images; and according to the real-time reconstruction scheme, performing depth estimation and three-dimensional scene reconstruction on the plurality of processed real-time two-dimensional images by using a three-dimensional scene reconstruction model to obtain real-time three-dimensional scene data. According to the method, the problems of high cost investment, low precision, low automation degree and unstable reconstruction result in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of three-dimensional scene reconstruction, and particularly relates to a three-dimensional scene reconstruction method and system based on monocular depth estimation. Background Art

[0002] A large-scale unbounded indoor environment generally refers to a very large indoor space without obvious partitions or boundaries. Such an environment usually has the spatial characteristics of vast space, no clear boundaries, and multi-functionality. With the popularization of digital technologies, how to digitally transform a large-scale unbounded indoor environment has become an important research direction in this field. With the rapid development of technologies such as virtual reality, augmented reality, and intelligent robots, three-dimensional scene reconstruction technology has shown broad application prospects in the field of digital transformation of large-scale unbounded indoor environments.

[0003] However, there are still many challenges in the three-dimensional reconstruction and virtualization of large-scale unbounded indoor environments in the prior art, including the following defects: 1) High cost investment: The prior art often relies on multi-camera or depth cameras for three-dimensional reconstruction, and the costs of these devices are relatively high, which restricts the popularization and application of the technology; 2) Insufficient accuracy: The accuracy of three-dimensional reconstruction using monocular depth estimation in the prior art is relatively low. Especially in areas with complex textures, large lighting changes, or no textures, the accuracy of depth estimation is severely affected; 3) Low degree of automation: In the prior art, links such as shooting point setting, image processing, and three-dimensional reconstruction often require manual intervention, and the degree of automation is relatively low; 4) Unstable reconstruction results: Due to the uncertainty of depth estimation, the three-dimensional reconstruction results of the prior art are often unstable and are easily affected by noise and outliers. Summary of the Invention

[0004] In order to solve the problems of high cost investment, insufficient accuracy, low degree of automation, and unstable reconstruction results existing in the prior art, the purpose of the present invention is to provide a three-dimensional scene reconstruction method and system based on monocular depth estimation.

[0005] The technical solution adopted by the present invention is as follows: A three-dimensional scene reconstruction method based on monocular depth estimation, comprising the following steps: Using artificial intelligence algorithms to construct a reconstruction plan generation model, a two-dimensional image processing model, and a three-dimensional scene reconstruction model; According to the real-time indoor layout data of the large-scale unbounded indoor environment, using the reconstruction plan generation model to generate a reconstruction plan and obtain a real-time reconstruction plan; According to the real-time reconstruction plan, using a monocular camera to collect a number of real-time two-dimensional images of the large-scale unbounded indoor environment from multiple perspectives; According to the real-time reconstruction scheme, using a two-dimensional image processing model, a number of real-time two-dimensional images are processed to obtain a number of processed real-time two-dimensional images; According to the real-time reconstruction scheme, using a three-dimensional scene reconstruction model, depth estimation and three-dimensional scene reconstruction are performed on a number of processed real-time two-dimensional images to obtain real-time three-dimensional scene data.

[0006] Furthermore, the real-time reconstruction scheme includes real-time two-dimensional image acquisition decision-making, real-time two-dimensional image processing effect requirement decision-making, and real-time three-dimensional scene data reconstruction effect requirement decision-making.

[0007] Furthermore, using artificial intelligence algorithms, a reconstruction scheme generation model, a two-dimensional image processing model, and a three-dimensional scene reconstruction model are constructed, including the following steps: Collect a number of historical indoor layout data, and according to the number of historical indoor layout data, use a swarm intelligence optimization algorithm to construct a reconstruction scheme generation model to obtain a number of historical reconstruction schemes; the historical reconstruction scheme includes historical two-dimensional image acquisition decision-making, historical two-dimensional image processing effect requirement decision-making, and historical three-dimensional scene reconstruction effect requirement decision-making; Collect a number of historical two-dimensional images, and according to the number of historical two-dimensional images and the historical two-dimensional image processing effects of a number of historical reconstruction schemes, use a reinforcement learning and image processing algorithm to construct a two-dimensional image processing model to obtain a number of processed historical two-dimensional images; According to the historical three-dimensional scene reconstruction effect requirement decision-making of a number of processed historical two-dimensional images and a number of historical reconstruction schemes, use a deep learning and three-dimensional reconstruction algorithm to construct a three-dimensional scene reconstruction model.

[0008] Furthermore, the reconstruction scheme generation model is constructed based on the LSTM-ISGA algorithm, and the reconstruction scheme generation model includes an indoor layout feature extraction module constructed based on the LSTM algorithm and a reconstruction scheme generation module constructed based on the ISGA algorithm connected in sequence. The reconstruction scheme generation module includes an initial solution generation module, an iterative optimization module, and an optimal solution analysis module connected in sequence.

[0009] Furthermore, the two-dimensional image processing model is constructed based on the CNN-MOGRPO-IPA algorithm, and the two-dimensional image processing model includes a two-dimensional image feature extraction module constructed based on the CNN algorithm, a processing strategy generation module constructed based on the MOGRPO algorithm, and a two-dimensional image processing algorithm library provided with a number of IPA algorithms.

[0010] Furthermore, the 3D scene reconstruction model is constructed based on the FPN-ResNet-3D-DBN algorithm, and the 3D scene reconstruction model includes a multi-scale fusion feature extraction module constructed based on the FPN algorithm, a depth estimation module constructed based on the ResNet algorithm, and a 3D scene reconstruction module constructed based on the 3D-DBN algorithm, which are connected in sequence.

[0011] Furthermore, according to the real-time indoor layout data of a large-scale unbounded indoor environment, a model for generating a reconstruction plan is used to generate a reconstruction plan, and a real-time reconstruction plan is obtained, including the following steps: Use the indoor layout feature extraction module of the model for generating a reconstruction plan to extract the real-time indoor layout features of the real-time indoor layout data of the large-scale unbounded indoor environment, and input them into the reconstruction plan generation module; Use the initial solution generation module of the reconstruction plan generation module to generate initial solutions according to the real-time indoor layout features, and obtain a number of initial solutions; the initial solutions correspond to the initial real-time reconstruction plans; Use the iterative optimization module of the reconstruction plan generation module to perform iterative optimization on a number of initial solutions to obtain the optimal solution with the best fitness value; Use the optimal solution analysis module of the reconstruction plan generation module to analyze the individual vectors of the optimal solution to obtain the optimal real-time reconstruction plan.

[0012] Furthermore, according to the real-time reconstruction plan, a 2D image processing model is used to process a number of real-time 2D images to obtain a number of processed real-time 2D images, including the following steps: Use the 2D image feature extraction module of the 2D image processing model to extract the real-time 2D image features of the real-time 2D images; Use the processing strategy generation module of the 2D image processing model to generate a processing strategy according to the decision on the real-time 2D image processing effect requirements of the real-time reconstruction plan and the real-time 2D image features, and obtain a real-time 2D image processing strategy; According to the real-time 2D image processing strategy, call the corresponding target IPA algorithm in the 2D image processing algorithm library, and process the corresponding real-time 2D image according to the target IPA algorithm to obtain the corresponding processed real-time 2D image; Traverse all real-time 2D images, and repeat the above steps to obtain the processed real-time 2D image corresponding to each real-time 2D image.

[0013] Furthermore, according to the real-time reconstruction plan, a 3D scene reconstruction model is used to perform depth estimation and 3D scene reconstruction on a number of processed real-time 2D images to obtain real-time 3D scene data, including the following steps: Make a decision on the real-time 3D scene data reconstruction effect requirements according to the real-time reconstruction scheme, adjust the model parameters of the 3D scene reconstruction model, and obtain the adjusted 3D scene reconstruction model; Use the adjusted multi-scale fusion feature extraction module of the adjusted 3D scene reconstruction model to extract the real-time multi-scale fusion features of each processed real-time 2D image; Use the adjusted depth estimation module of the adjusted 3D scene reconstruction model to perform depth estimation according to each real-time multi-scale fusion feature, and obtain the corresponding real-time depth map; Use the adjusted 3D scene reconstruction module of the adjusted 3D scene reconstruction model to perform 3D scene reconstruction according to all real-time depth maps, and obtain real-time 3D scene data.

[0014] A 3D scene reconstruction system based on monocular depth estimation is used to implement the 3D scene reconstruction method. The system includes a model construction unit, a reconstruction scheme generation unit, a 2D image processing unit, and a depth estimation and 3D scene reconstruction unit connected in sequence.

[0015] The beneficial effects of the present invention are as follows: The present invention discloses a 3D scene reconstruction method and system based on monocular depth estimation. By using a monocular camera to replace a multi-camera or a depth camera, the equipment cost investment is greatly reduced, making the 3D reconstruction technology more affordable; the artificial intelligence algorithm is used to optimize the 2D feature extraction and depth estimation processes, effectively improving the accuracy of monocular depth estimation, especially in complex texture and textureless areas; through the constructed reconstruction scheme generation model, 2D image processing model, and 3D scene reconstruction model, automatic reconstruction scheme generation, 2D image processing, and 3D scene reconstruction are realized, achieving full automation from shooting point setting to 3D reconstruction, reducing manual intervention, improving efficiency, adapting to reconstruction schemes for large-scale unbounded indoor environment changes, and enhancing the versatility and flexibility of the technology; the 2D image processing model generated by the adopted adaptive processing strategy and the 3D scene reconstruction model optimized by multi-scale fusion feature extraction and depth estimation can maintain high robustness and accuracy under variable lighting and complex texture conditions, improving the stability and reliability of the 3D reconstruction results.

[0016] Other beneficial effects of the present invention will be further described in the specific implementation manner. Brief Description of the Drawings

[0017] Figure 1 is a flowchart of the 3D scene reconstruction method based on monocular depth estimation in the present invention.

[0018] Figure 2 is a structural block diagram of the 3D scene reconstruction system based on monocular depth estimation in the present invention. Detailed implementation manners

[0019] The present invention will be further explained below with reference to the accompanying drawings and specific embodiments.

[0020] Embodiment 1: As Figure 1 shown, this embodiment provides a three-dimensional scene reconstruction method based on monocular depth estimation, including the following steps: S1: Using artificial intelligence algorithms, construct a reconstruction plan generation model, a two-dimensional image processing model, and a three-dimensional scene reconstruction model, including the following steps: S1-1: Collect a number of historical indoor layout data, and based on the number of historical indoor layout data, use a swarm intelligence optimization algorithm to construct a reconstruction plan generation model to obtain a number of historical reconstruction plans; the historical reconstruction plans include historical two-dimensional image acquisition decision-making, historical two-dimensional image processing effect requirement decision-making, and historical three-dimensional scene reconstruction effect requirement decision-making; The reconstruction plan generation model is constructed based on the Long Short-Term Memory (LSTM)-Improved Snow Geese Algorithm (ISGA), and the reconstruction plan generation model includes an indoor layout feature extraction module constructed based on the LSTM algorithm and a reconstruction plan generation module constructed based on the ISGA algorithm connected in sequence. The reconstruction plan generation module includes an initial solution generation module, an iterative optimization module, and an optimal solution analysis module connected in sequence; LSTM can capture the spatio-temporal dependence relationship in the indoor layout data, effectively extract the key features of the indoor layout, including size, shape, layout position, etc., and effectively retain the spatio-temporal information in the indoor layout data, providing accurate basic features for the subsequent generation of reconstruction plans; the ISGA algorithm has a powerful global search ability, can find a reconstruction plan close to the optimal one, and can automatically adjust the reconstruction plan according to different indoor layouts, with strong adaptability; S1-2: Collect a number of historical two-dimensional images, and based on the historical two-dimensional image processing effect requirement decision-making of the number of historical two-dimensional images and the number of historical reconstruction plans, use a reinforcement learning and image processing algorithm to construct a two-dimensional image processing model to obtain a number of processed historical two-dimensional images; The two-dimensional image processing model is constructed based on the Convolutional Neural Network (CNN) - Multi-Objective Group Relative Policy Optimization (MOGRPO) - Image Processing Algorithm (IPA). The two-dimensional image processing model includes a two-dimensional image feature extraction module constructed based on the CNN algorithm, a processing strategy generation module constructed based on the MOGRPO algorithm, and a two-dimensional image processing algorithm library with several IPA algorithms connected in sequence; The two-dimensional image feature extraction module captures low-level features such as edges, textures, and shapes in the image through operations such as multi-layer convolution and pooling. It is efficient and can quickly and accurately identify the key information in the image as the basis for generating processing strategies. The processing strategy generation module can consider multiple optimization objectives simultaneously, such as processing speed, accuracy, and robustness, and generate a comprehensive optimal processing strategy. It can automatically adjust and optimize the processing strategy according to different image characteristics and requirements, with strong adaptability. The two-dimensional image processing algorithm library provides a variety of image processing algorithms, such as filtering processing, enhancement processing, segmentation processing, morphological processing, and perspective alignment processing, etc., which can meet the processing requirements of different scenarios and needs. It can be flexibly configured and select appropriate image processing algorithms to achieve personalized processing and provide high-quality image data for subsequent three-dimensional scene reconstruction; The processing strategy generation module includes an objective function set, a policy network, an experience replay pool, and an agent. The agent is respectively connected to the objective function set and the policy network. The agent learns historical processing strategies through the experience replay pool and continuously optimizes its own strategy generation ability. The agent controls the policy network according to the learned experience to generate more effective processing strategies. The design of the experience replay pool and the agent enables the model to continuously learn and optimize, improving the quality of strategy generation. Since the processing strategy generation module adopts a group exploration method, it can avoid falling into local optimal solutions to a certain extent. The policy network outputs the distribution probability of actions in a given state. The processing strategy generation module directly updates the policy network through gradients, eliminating the Critic model in traditional reinforcement learning and making the algorithm structure more concise; S1-3: Make a decision based on the historical three-dimensional scene reconstruction effect requirements of several processed historical two-dimensional images and several historical reconstruction schemes, and use deep learning and three-dimensional reconstruction algorithms to construct a three-dimensional scene reconstruction model; The three-dimensional scene reconstruction model is constructed based on the Feature Pyramid Network (FPN)-Residual Network (ResNet)-Three-Dimensional Deep Belief Network (3D-DBN) algorithm. The three-dimensional scene reconstruction model includes a multi-scale fusion feature extraction module constructed based on the FPN algorithm, a depth estimation module constructed based on the ResNet algorithm, and a three-dimensional scene reconstruction module constructed based on the 3D-DBN algorithm, which are connected in sequence; The FPN of the multi-scale fusion feature extraction module combines the top-down path and the bottom-up path to achieve the fusion of features at different scales, captures image information from fine-grained to coarse-grained, effectively integrates image features at different scales, enhances the ability of feature representation, is beneficial to the accuracy of subsequent depth estimation, retains the detailed information of the image, and fuses the global semantic information, providing a more comprehensive feature basis for depth estimation. The fusion of multi-scale features helps to improve the adaptability of the depth estimation module to complex scenes and reduce the estimation error caused by scale changes; the ResNet of the depth estimation module solves the problem of gradient disappearance in the training of deep neural networks by introducing the residual learning mechanism, achieves more accurate depth estimation, helps to speed up the network training speed, reduces the training time, and has stronger robustness to complex scenes with different illuminations, textures, etc., ensuring the stability of depth estimation; the three-dimensional scene reconstruction module is trained with a large number of historical depth maps to learn the prediction from depth maps to three-dimensional scene data. 3D-DBN is good at processing three-dimensional data, can generate high-quality three-dimensional scene models, restore the three-dimensional structure of the scene, and can better restore the detailed information in the scene, such as the shape and position relationship of objects, and is applicable to the three-dimensional scene reconstruction of various complex large-scale unbounded indoor environments; S2: According to the real-time indoor layout data of the large-scale unbounded indoor environment, use the reconstruction scheme generation model to generate a model and perform reconstruction scheme generation to obtain a real-time reconstruction scheme, including the following steps: S2-1: Use the indoor layout feature extraction module of the reconstruction scheme generation model to extract the real-time indoor layout features of the real-time indoor layout data of the large-scale unbounded indoor environment and input them into the reconstruction scheme generation module; S2-2: Use the initial solution generation module of the reconstruction scheme generation module to generate initial solutions according to the real-time indoor layout features to obtain a number of initial solutions; the initial solutions correspond to the initial real-time reconstruction scheme; The formula is:

[0021] In the formula, The initial ISGA individuals generated for the Circle chaotic mapping sequence, i.e., the initial solutions; ISGA individuals randomly generated; i The ISGA individual indicator; Using the Circle chaotic mapping sequence to generate the initial population, compared with the randomly distributed population, the initial position distribution of the improved ISGA population is more uniform, expanding the search range of the ISGA population in space, increasing the diversity of the population positions, and improving to a certain extent the defect that the algorithm is prone to fall into local extrema, thus improving the optimization efficiency of the algorithm; S2-3: Use the reconstruction scheme to generate the iterative optimization module of the module, and perform iterative optimization on several initial solutions to obtain the optimal solution with the best fitness value, including the following steps: S2-3-1: Take several initial solutions as the initial ISGA population; Each initial ISGA individual in the initial ISGA population corresponds to an initial solution; S2-3-2: Use the fitness function and the iterative update module to obtain the initial fitness value of each initial ISGA individual in the initial ISGA population, and take the initial ISGA individual with the lowest fitness value as the leading goose; The formula is:

[0022] In the formula, is the real-time fitness function; is the two-dimensional image acquisition cost function; is the two-dimensional image processing function; is the three-dimensional scene data reconstruction cost; is the ISGA individual; are the first weight value, the second weight value, and the third weight value; S2-3-3: Enter the exploration stage, introduce the leading goose rotation mechanism, the calling guidance mechanism, and the dynamic reverse mechanism, perform iterative update on the initial ISGA population to obtain the once-updated ISGA population, and retain the optimal individual; The leading goose rotation mechanism, in each iteration, competes according to the fitness value of the ISGA individual to select a new leading goose. This mechanism can avoid the leading goose falling into local optimum prematurely and enhance the global search ability of the algorithm; The formula is:

[0023] In the formula, is the once-updated leading goose; is the initial ISGA individual with the third-lowest reciprocal fitness value in the initial ISGA population at the iteration number; The initial ISGA individual ranked fifth from the bottom in terms of fitness value in the initial ISGA population of the iteration number; Is the current iteration number; Is the optimal individual; Is the first weight factor; Is a random number; Is the initial leading goose; The calling guidance mechanism adjusts the individual position update using the acoustic wave propagation attenuation model according to the distance between the ISGA individual and the leading goose. For the ISGA individuals with a relatively close distance, their position updates are more affected by the leading goose and can quickly approach the optimal solution. For the ISGA individuals with a relatively far distance, their position updates are less affected by the leading goose and can maintain a certain exploration ability. This mechanism can avoid excessive aggregation or dispersion of the group and improve the local search accuracy of the algorithm; The formula is:

[0024] In the formula, Is the ISGA individual updated once; Is the Initial ISGA individual of the iteration number; Is the initial ISGA individual Received sound intensity; Is the Initial ISGA individual of the iteration number Corresponding sound intensity parameter; Is the initial sound intensity; Is the minimum receivable sound intensity; Is the convergence factor; Is the initial ISGA individual with the farthest distance; Is the random parameter; Is the Brownian motion function; Is the Brownian motion parameter; Is the exclusive OR processing symbol;

[0025] In the formula, Is the convergence factor; tanh(.) is the hyperbolic tangent function; Is the current iteration number; Is the maximum iteration number; a max 、 a min Are the maximum and minimum values of the convergence factor respectively; λ Is the decreasing rate parameter, Is the decreasing period parameter, λ =-2 π , = π ; A dynamic reverse mechanism is used to dynamically reverse the initial ISGA individuals, improve the diversity of exploration directions, and avoid falling into local optima; The formula is:

[0026] In the formula, is the reverse ISGA individual updated once; γ is the decreasing inertia coefficient; L max and L min are the maximum and minimum values of the vector space respectively; The leading goose updated once, several ISGA individuals updated once, and several reverse ISGA individuals updated once will be integrated to obtain the ISGA population updated once, and the ISGA individual with the lowest fitness value will be retained as the optimal individual; S2-3-4: Enter the development stage, introduce the abnormal boundary strategy and the Gaussian mutation mechanism, perform a second update on the ISGA population updated once to obtain the ISGA population updated twice, and retain the optimal individual; For the abnormal boundary strategy, calculate the difference between the fitness value of each ISGA individual updated once and the average fitness value of the population. For the ISGA individuals with fitness values much higher than the average value of the population, their position update methods will be adjusted, such as using the Gaussian mutation mechanism, a larger step size or a smaller step size. This mechanism can help individuals avoid falling into local optima and improve the convergence speed and accuracy of the algorithm; The formula is:

[0027] In the formula, is the ISGA individual updated twice; is the ISGA individual updated once; is the fitness function; is the average fitness value of the population; is the ISGA individual with the highest fitness value; are the second weight factor and the third weight factor; is the parameter of the Gaussian mutation mechanism; S2-3-5: If the number of iterations is greater than or equal to the iteration number threshold or the fitness value of the optimal individual is less than the fitness threshold, then output the optimal individual as the optimal solution; S2-4: Use the reconstruction scheme to generate the optimal solution analysis module of the module, analyze the individual vector of the optimal solution, and obtain the optimal real-time reconstruction scheme; The real-time reconstruction scheme includes real-time two-dimensional image acquisition decision-making, real-time two-dimensional image processing effect requirement decision-making, and real-time three-dimensional scene data reconstruction effect requirement decision-making; The real-time two-dimensional image acquisition decision-making includes the number of real-time monocular cameras, the number of real-time shooting points, the pose of real-time monocular cameras, the shooting time of real-time monocular cameras, etc.; The real-time two-dimensional image processing effect requirement decision-making includes the real-time image quality requirements, real-time image viewing angle requirements, real-time image size requirements, etc. for the processed real-time two-dimensional images; The real-time three-dimensional scene data reconstruction effect requirement decision-making includes the real-time data type requirements, real-time granularity requirements, real-time accuracy requirements, etc. for the real-time three-dimensional scene data; S3: According to the real-time reconstruction scheme, use monocular cameras to collect a number of real-time two-dimensional images of a large-scale unbounded indoor environment from multiple perspectives; S4: According to the real-time reconstruction scheme, use a two-dimensional image processing model to process a number of real-time two-dimensional images to obtain a number of processed real-time two-dimensional images, including the following steps: S4-1: Use the two-dimensional image feature extraction module of the two-dimensional image processing model to extract the real-time two-dimensional image features of the real-time two-dimensional images; S4-2: Use the processing strategy generation module of the two-dimensional image processing model to generate a processing strategy based on the real-time two-dimensional image processing effect requirement decision-making and real-time two-dimensional image features of the real-time reconstruction scheme to obtain a real-time two-dimensional image processing strategy; S4-3: According to the real-time two-dimensional image processing strategy, call the corresponding target IPA algorithm in the two-dimensional image processing algorithm library and process the corresponding real-time two-dimensional image according to the target IPA algorithm to obtain the corresponding processed real-time two-dimensional image; S4-4: Traverse all real-time two-dimensional images and repeat the above steps to obtain the processed real-time two-dimensional image corresponding to each real-time two-dimensional image; S5: According to the real-time reconstruction scheme, use a three-dimensional scene reconstruction model to perform depth estimation and three-dimensional scene reconstruction on a number of processed real-time two-dimensional images to obtain real-time three-dimensional scene data, including the following steps: S5-1: According to the real-time three-dimensional scene data reconstruction effect requirement decision-making of the real-time reconstruction scheme, adjust the model parameters of the three-dimensional scene reconstruction model to obtain an adjusted three-dimensional scene reconstruction model; S5-2: Use the adjusted multi-scale fusion feature extraction module of the adjusted three-dimensional scene reconstruction model to extract the real-time multi-scale fusion features of each processed real-time two-dimensional image; S5-3: Use the adjusted depth estimation module of the adjusted three-dimensional scene reconstruction model to perform depth estimation according to each real-time multi-scale fusion feature to obtain the corresponding real-time depth map; S5-4: The adjusted 3D scene reconstruction module using the adjusted 3D scene reconstruction model performs 3D scene reconstruction based on all real-time depth maps to obtain real-time 3D scene data.

[0028] Embodiment 2: As Figure 2 shown, this embodiment provides a 3D scene reconstruction system based on monocular depth estimation for implementing a 3D scene reconstruction method. The system includes a model construction unit, a reconstruction scheme generation unit, a 2D image processing unit, and a depth estimation and 3D scene reconstruction unit connected in sequence; The model construction unit is used to construct a reconstruction scheme generation model, a 2D image processing model, and a 3D scene reconstruction model using artificial intelligence algorithms; The reconstruction scheme generation unit is used to generate a real-time reconstruction scheme according to the real-time indoor layout data of a large-scale unbounded indoor environment using the reconstruction scheme generation model; The 2D image processing unit is used to process a number of real-time 2D images collected by a monocular camera using the 2D image processing model according to the real-time reconstruction scheme to obtain a number of processed real-time 2D images; The depth estimation and 3D scene reconstruction unit is used to perform depth estimation and 3D scene reconstruction on a number of processed real-time 2D images using the 3D scene reconstruction model according to the real-time reconstruction scheme to obtain real-time 3D scene data.

[0029] The present invention discloses a 3D scene reconstruction method and system based on monocular depth estimation. By using a monocular camera to replace a multi-camera or depth camera, the equipment cost investment is greatly reduced, making the 3D reconstruction technology more affordable; the artificial intelligence algorithm is used to optimize the 2D feature extraction and depth estimation processes, effectively improving the accuracy of monocular depth estimation, especially in complex texture and textureless areas; through the constructed reconstruction scheme generation model, 2D image processing model, and 3D scene reconstruction model, automated reconstruction scheme generation, 2D image processing, and 3D scene reconstruction are realized, achieving full automation from shooting point setting to 3D reconstruction, reducing manual intervention, improving efficiency, adapting to reconstruction schemes for large-scale unbounded indoor environment changes, and enhancing the versatility and flexibility of the technology; the 2D image processing model generated by the adopted adaptive processing strategy and the 3D scene reconstruction model with multi-scale fusion feature extraction and depth estimation optimization can maintain high robustness and accuracy under variable lighting and complex texture conditions, improving the stability and reliability of the 3D reconstruction results.

[0030] The present invention is not limited to the above optional embodiments, and any person can obtain other various forms of products under the inspiration of the present invention. The above specific embodiments should not be construed as limiting the protection scope of the present invention. The protection scope of the present invention shall be defined by the claims, and the specification can be used to interpret the claims.

Claims

1. A three-dimensional scene reconstruction method based on monocular depth estimation, characterized in that: The steps include: Use artificial intelligence algorithms to build reconstruction solution generation models, 2D image processing models, and 3D scene reconstruction models; According to the real-time indoor layout data of a large-scale unbounded indoor environment, a reconstruction scheme generation model is used to generate a reconstruction scheme to obtain a real-time reconstruction scheme; According to the real-time reconstruction scheme, a monocular camera is used to collect several real-time two-dimensional images of a large-scale unbounded indoor environment from multiple perspectives; According to the real-time reconstruction scheme, a two-dimensional image processing model is used to process a plurality of real-time two-dimensional images to obtain a plurality of processed real-time two-dimensional images; According to the real-time reconstruction scheme, a three-dimensional scene reconstruction model is used to perform depth estimation and three-dimensional scene reconstruction on a number of processed real-time two-dimensional images to obtain real-time three-dimensional scene data.

2. The method for 3D scene reconstruction based on monocular depth estimation according to claim 1, characterized in that: The real-time reconstruction scheme includes real-time two-dimensional image acquisition decision, real-time two-dimensional image processing effect requirement decision and real-time three-dimensional scene data reconstruction effect requirement decision.

3. The method for 3D scene reconstruction based on monocular depth estimation according to claim 2, characterized in that: Using artificial intelligence algorithms, we build a reconstruction solution generation model, a two-dimensional image processing model, and a three-dimensional scene reconstruction model, including the following steps: Collecting a number of historical indoor layout data, and using a swarm intelligence optimization algorithm based on the number of historical indoor layout data to construct a reconstruction scheme generation model to obtain a number of historical reconstruction schemes; the historical reconstruction schemes include historical two-dimensional image acquisition decisions, historical two-dimensional image processing effect requirements decisions, and historical three-dimensional scene reconstruction effect requirements decisions; Collect a number of historical two-dimensional images, and use reinforcement learning and image processing algorithms to construct a two-dimensional image processing model based on the historical two-dimensional images and the historical two-dimensional image processing effects of the historical reconstruction schemes to obtain a number of processed historical two-dimensional images; Based on the historical three-dimensional scene reconstruction effect requirements of several processed historical two-dimensional images and several historical reconstruction schemes, a three-dimensional scene reconstruction model is constructed using deep learning and three-dimensional reconstruction algorithms.

4. The method for 3D scene reconstruction based on monocular depth estimation according to claim 3, characterized in that: The reconstruction scheme generation model is constructed based on the LSTM-ISGA algorithm, and the reconstruction scheme generation model includes an indoor layout feature extraction module constructed based on the LSTM algorithm and a reconstruction scheme generation module constructed based on the ISGA algorithm, which are connected in sequence. The reconstruction scheme generation module includes an initial solution generation module, an iterative optimization module and an optimal solution analysis module, which are connected in sequence.

5. The method for 3D scene reconstruction based on monocular depth estimation according to claim 4, characterized in that: The two-dimensional image processing model is constructed based on the CNN-MOGRPO-IPA algorithm, and the two-dimensional image processing model includes a two-dimensional image feature extraction module constructed based on the CNN algorithm, a processing strategy generation module constructed based on the MOGRPO algorithm, and a two-dimensional image processing algorithm library provided with several IPA algorithms, which are connected in sequence.

6. The method for 3D scene reconstruction based on monocular depth estimation according to claim 5, characterized in that: The three-dimensional scene reconstruction model is constructed based on the FPN-ResNet-3D-DBN algorithm, and the three-dimensional scene reconstruction model includes a multi-scale fusion feature extraction module constructed based on the FPN algorithm, a depth estimation module constructed based on the ResNet algorithm, and a three-dimensional scene reconstruction module constructed based on the 3D-DBN algorithm, which are connected in sequence.

7. The method for 3D scene reconstruction based on monocular depth estimation according to claim 6, characterized in that: According to the real-time indoor layout data of a large-scale unbounded indoor environment, a reconstruction scheme generation model is used to generate a reconstruction scheme to obtain a real-time reconstruction scheme, including the following steps: Using the indoor layout feature extraction module of the reconstruction solution generation model, the real-time indoor layout features of the real-time indoor layout data of the large-scale unbounded indoor environment are extracted and input into the reconstruction solution generation module; Using the initial solution generation module of the reconstruction solution generation module, an initial solution is generated according to the real-time indoor layout characteristics to obtain a number of initial solutions; the initial solutions correspond to the initial real-time reconstruction solutions; Use the iterative optimization module of the reconstruction solution generation module to iteratively optimize several initial solutions to obtain the optimal solution with the best fitness value; The optimal solution analysis module of the reconstruction solution generation module is used to analyze the individual vectors of the optimal solution to obtain the optimal real-time reconstruction solution.

8. The method for 3D scene reconstruction based on monocular depth estimation according to claim 7, characterized in that: According to the real-time reconstruction scheme, a two-dimensional image processing model is used to process a plurality of real-time two-dimensional images to obtain a plurality of processed real-time two-dimensional images, including the following steps: Using a two-dimensional image feature extraction module of a two-dimensional image processing model to extract real-time two-dimensional image features of a real-time two-dimensional image; Using a processing strategy generation module of a two-dimensional image processing model, a processing strategy is generated according to a real-time two-dimensional image processing effect requirement decision of a real-time reconstruction scheme and real-time two-dimensional image features to obtain a real-time two-dimensional image processing strategy; According to the real-time two-dimensional image processing strategy, the corresponding target IPA algorithm is called in the two-dimensional image processing algorithm library, and the corresponding real-time two-dimensional image is processed according to the target IPA algorithm to obtain the corresponding processed real-time two-dimensional image; Traverse all real-time two-dimensional images and repeat the above steps to obtain a processed real-time two-dimensional image corresponding to each real-time two-dimensional image.

9. The method for 3D scene reconstruction based on monocular depth estimation according to claim 8, characterized in that: According to the real-time reconstruction scheme, a 3D scene reconstruction model is used to perform depth estimation and 3D scene reconstruction on a number of processed real-time 2D images to obtain real-time 3D scene data, including the following steps: According to the real-time three-dimensional scene data reconstruction effect demand decision of the real-time reconstruction solution, the model parameters of the three-dimensional scene reconstruction model are adjusted to obtain an adjusted three-dimensional scene reconstruction model; Extracting real-time multi-scale fusion features of each processed real-time two-dimensional image using the adjusted multi-scale fusion feature extraction module of the adjusted three-dimensional scene reconstruction model; Using the adjusted depth estimation module of the adjusted 3D scene reconstruction model, performing depth estimation according to each real-time multi-scale fusion feature to obtain a corresponding real-time depth map; An adjusted 3D scene reconstruction module of the adjusted 3D scene reconstruction model is used to reconstruct the 3D scene according to all real-time depth maps to obtain real-time 3D scene data.

10. A three-dimensional scene reconstruction system based on monocular depth estimation, used to implement the three-dimensional scene reconstruction method according to any one of claims 1 to 9, characterized in that: The system comprises a model building unit, a reconstruction scheme generating unit, a two-dimensional image processing unit and a depth estimation and three-dimensional scene reconstruction unit which are connected in sequence.

Citation Information

Patent Citations

  • Three-dimensional reconstruction method based on attention mechanism and monocular multi-view angle

    CN113838191A

  • Three-dimensional scene reconstruction method and system based on mixed reality glasses and application

    CN114219900A

  • Natural landscape multi-view three-dimensional reconstruction method based on deep learning

    CN114677479A

  • Three-dimensional scene reconstruction method and device based on neural network and multi-view consistency

    CN117523100A

  • Light field three-dimensional imaging method and system based on neural network

    CN117934708A

Cited By

  • Three-dimensional scene reconstruction method and system based on monocular depth estimation

    CN121482285A

  • Three-dimensional scene reconstruction method and system based on monocular depth estimation

    CN121482285B