Scene building method and system based on three-dimensional reconstruction
Through the scene construction method based on three-dimensional reconstruction, multi-view image acquisition and semantic segmentation technology, combined with beam method adjustment algorithm, the technical bottlenecks of traditional technology when rendering large-scale complex scenes in real-time, achieving efficient and accurate scene construction and automated combination.
Patent Information
- Application Number
- CN202510398756.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-01
AI Technical Summary
The existing technology has significant technical bottlenecks when rendering large-scale complex scenes in real time. The traditional three-dimensional scene construction method is time-consuming and labor-intensive, and it is difficult to meet the requirements of efficiency and accuracy.
Using a scene construction method based on three-dimensional reconstruction, a three-dimensional reconstruction algorithm based on multi-view image acquisition, preprocessing, semantic segmentation and beam method adjustment, the objects and regions in the image are automatically identified and marked, and a high-precision three-dimensional model is reconstructed.
It effectively solves the technical bottleneck of real-time rendering of large-scale complex scenes, improves the efficiency and accuracy of scene construction, guarantees user experience, and realizes the automation and intelligence of scene construction.
Smart Images

Figure CN119991970A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a scene building method and system based on three-dimensional reconstruction. Background Art
[0002] With the development of computer graphics, virtual reality (VR) and augmented reality (AR) technologies have been widely used in entertainment, education, medical and other fields. However, existing technologies still have significant technical bottlenecks in real-time rendering of large-scale complex scenes. Traditional 3D scene construction methods usually rely on professional 3D modeling software, requiring designers to manually create 3D models and perform texture mapping, which is not only time-consuming and labor-intensive, but also difficult to meet the efficiency and accuracy requirements of large-scale scenes in real-time rendering, resulting in the user experience in complex scenes cannot be fully guaranteed. Summary of the invention
[0003] The purpose of the present invention is to provide a scene construction method and system based on three-dimensional reconstruction, which effectively solves the technical bottleneck of real-time rendering of large-scale complex scenes and improves the efficiency and accuracy of scene construction through multi-view image acquisition, preprocessing, semantic segmentation and a three-dimensional reconstruction algorithm based on bundle adjustment.
[0004] The purpose of the present invention can be achieved through the following technical solutions: This application provides a scene construction method based on three-dimensional reconstruction, comprising the following steps: Collect multi-view 2D images of the environment through drones, mobile phone cameras, and image acquisition devices; Preprocess the acquired two-dimensional images and use semantic segmentation technology to automatically identify and annotate different objects and areas in the images and understand the semantic information in the images; According to the annotated two-dimensional image data, a three-dimensional model is reconstructed from the two-dimensional image using a multi-viewpoint geometric three-dimensional reconstruction algorithm based on bundle adjustment, and the reconstructed three-dimensional model is stored as a sub-scene model; The method utilizes a multi-viewpoint geometric 3D reconstruction algorithm based on bundle adjustment to reconstruct a 3D model from a 2D image and stores the reconstructed 3D model as a subscene model, specifically comprising: Match the corresponding feature points between images of different perspectives, determine the positions of the feature points in different images, and determine the intrinsic and extrinsic parameters of the camera through the known calibration plate image; Using multi-view feature point matching and camera parameters, the three-dimensional coordinates of each feature point in space are determined by triangulation method; Optimize the 3D coordinates and camera parameters by minimizing the reprojection error, which refers to the difference between the 3D point projected back to the 2D image and the actual feature point position, and use the least squares method to update the parameter estimates; The minimized reprojection error is expressed as: ; in, is the location of the feature point in the image, is the corresponding 3D point position, are the camera parameters, represents the projection function, is a robust kernel function used to handle outliers; According to the optimized coordinates of the 3D points, the camera parameters of the image are stored as a sub-scene model.
[0005] Furthermore, the acquired two-dimensional image is preprocessed, specifically including: Remove noise from 2D images by bilateral filtering; Specifically expressed as: ; in, is the pixel value of the filtered image at point (x,y), is the original image at point The pixel value of is a spatial Gaussian function, is the spatial standard deviation, controlling the size of the filter, is the intensity Gaussian function, is the intensity standard deviation, controlling the similarity between pixel values, is the normalization factor; The grayscale world hypothesis algorithm is used to adjust the gain of each channel for color correction, and then the white balance algorithm is used to adjust the color according to the grayscale card in the scene or according to the information of the image itself; The specific gray world hypothesis algorithm is expressed as: ;in is the average brightness of the entire image, is the average brightness of channel c; The specific white balance algorithm is expressed as: ; where I(c) is a channel of the original image, is the value of this channel after white balancing; The acquired two-dimensional images are registered by feature matching algorithm and image registration algorithm; Specifically, a feature detection algorithm is used to detect key points and descriptors, and then matching point pairs are found by matching descriptors; a transformation matrix for aligning one image to another image is solved by a transformation matrix; the transformation matrix solution includes a homography matrix and an affine transformation; The homography matrix is expressed as: ; where x and is a pair of matching points, is a robust kernel function; The affine transformation is expressed as: ;in and are the coordinates of the matching point pairs.
[0006] Furthermore, the use of semantic segmentation technology to automatically identify and label different objects and regions in the image specifically includes: Extract features from the preprocessed image through convolutional and pooling layers, and use a fully convolutional network for pixel-level classification; Use the softmax algorithm to classify each pixel, and use upsampling and deconvolution techniques to enlarge the feature map to its original size to generate a segmentation map; Through threshold processing and connected component analysis, pixel labels are aggregated into coherent object regions, and the boundaries of different objects are identified and located.
[0007] Furthermore, the use of the least squares method to update the parameter estimates specifically includes: ; where X is the design matrix, Y is the observation vector, is the parameter vector.
[0008] Furthermore, a scene construction method based on three-dimensional reconstruction further includes: automatically selecting appropriate stored sub-scene models by specifying the required scene type and specific requirements, and performing preliminary combination according to the automatically selected instructions to form an initial scene model; Among them, the stored sub-scene models are indexed by defining keywords and tags, and according to the results of project demand analysis, intelligent retrieval is used to perform keyword matching and tag screening to find sub-scene models that meet the requirements; then computer graphics technology is used to perform spatial analysis and calculation on the selected sub-scene models to determine the relative position and orientation of the sub-scene models in the virtual scene, and the models are automatically arranged and combined through a constraint-based layout algorithm.
[0009] Furthermore, a scene building method based on three-dimensional reconstruction also includes: automatically making minor adjustments to the initial scene model to generate a target scene description file of the target scene model, wherein the target scene description file includes information for rendering and displaying the scene.
[0010] Furthermore, the initial scene model is automatically fine-tuned by adding or removing objects, modifying materials, and adjusting lighting for training. Specifically, computer graphics techniques are used to automatically add or remove objects from a scene; Modify the material parameters by adjusting parameters such as reflectivity in the Blinn-Phong lighting model; adjust the material properties of the object surface, such as color, texture, roughness, etc., to match the style of the target scene; Use lighting models, a physically based rendering technique, to simulate real-world lighting effects and optimize lighting conditions in the scene, including adding or adjusting the position, intensity, and type of light sources.
[0011] In the physical rendering technology, the illumination calculation using the BRDF illumination model is specifically expressed as: ; in, is the Fresnel coefficient, which represents the reflectivity of the material. is the normal distribution function, which describes the effect of surface roughness on lighting. is a geometric function that describes the degree to which light is blocked by a surface. is the surface roughness, N is the surface normal, L is the direction vector from the surface point to the light source, V is the direction vector from the surface point to the camera, is the intensity of the light source.
[0012] Furthermore, after the initial scene model is automatically fine-tuned to generate a target scene description file of the target scene model, the target scene model generated by optimizing the generated scene model by combining the generative adversarial network and the variational autoencoder includes: fusing the loss function of the variational autoencoder and the loss function of the generative adversarial network to obtain a joint optimization target for obtaining a three-dimensional model; The joint loss function is expressed as: ; in, and is the weight coefficient that balances the two loss terms, Expressed as reconstruction loss, it is used to measure the similarity between the 3D point cloud reconstructed by the latent variable z and the original point cloud; Expressed as KL divergence, it is used to measure the distribution of latent variables With prior distribution The difference between Indicates that it encourages the discriminator to correctly identify the real 3D model. This is expressed as encouraging the generator to produce realistic fake 3D models.
[0013] Furthermore, a scene building method based on three-dimensional reconstruction also includes: based on the target scene description file for generating the target scene model, checking whether the generated target scene description file complies with preset rules, and deploying the model that complies with the rules to the virtual platform.
[0014] A scene building system based on three-dimensional reconstruction, including an image acquisition module, an image preprocessing module, a feature extraction and image registration module, a semantic segmentation module, a three-dimensional reconstruction and optimization module, and a scene combination and adjustment module; The image acquisition module acquires multi-view two-dimensional images of the environment through drones, mobile phone cameras and image acquisition devices; The image preprocessing module performs preprocessing operations on the collected two-dimensional image, specifically including noise removal by bilateral filtering technology, and color correction by grayscale world hypothesis and white balance algorithm; The feature extraction and image registration module uses a feature detection algorithm to detect key points and descriptors, and finds matching point pairs through a feature matching algorithm; Among them, the multi-view images are accurately aligned through the image registration algorithm of homography matrix and affine transformation; The semantic segmentation module uses the deep learning technology of the fully convolutional network to perform pixel-level semantic segmentation, automatically identify and annotate different objects and regions in the image, and understand the semantic information in the image; The three-dimensional reconstruction and optimization module reconstructs a three-dimensional model from a two-dimensional image based on a multi-viewpoint geometric three-dimensional reconstruction algorithm of bundle adjustment, and stores the reconstructed three-dimensional model as a sub-scene model; The scene combination and adjustment module automatically selects appropriate sub-scene models through intelligent retrieval technology, performs preliminary combination to form an initial scene model, and then makes minor adjustments to the initial scene model, including adding or removing objects, modifying materials and adjusting lighting, generating a target scene description file, and then deploying the model that meets the rules to the virtual platform; This also includes optimizing the initial scene model using generative adversarial networks and variational autoencoders.
[0015] The beneficial effects of the present invention are: (1) The present invention collects multi-view 2D images of the environment, pre-processes the acquired 2D images, and then uses advanced semantic segmentation technology to automatically identify and annotate different objects and areas in the images. Then, based on the annotated 2D image data, a multi-viewpoint geometric 3D reconstruction algorithm based on bundle adjustment is used to accurately reconstruct a 3D model from the 2D images. This effectively solves the technical bottlenecks of traditional 3D scene construction methods in real-time rendering of large-scale complex scenes, such as large amount of calculation and low precision. By improving the efficiency and precision of scene construction, the user experience is fully guaranteed. (2) The user only needs to specify the required scene type and specific requirements, automatically select the appropriate stored sub-scene model, and perform preliminary combination according to the automatically selected instructions to form an initial scene model. The initial scene model can be fine-tuned to achieve automation and intelligence in scene construction, greatly improving the speed and efficiency of scene construction. At the same time, since the need for manual intervention is reduced, the cost and error rate are also reduced. (3) Generative adversarial networks and variational autoencoders are combined to optimize the initial scene model. By integrating the joint optimization objectives, a more realistic and detailed three-dimensional model is obtained. This not only improves the authenticity and detail expression of the three-dimensional model, but also enhances the diversity and flexibility of the model. By combining GAN and VAE, higher quality and more realistic three-dimensional models can be generated, providing richer and more realistic scene experiences for fields such as virtual reality and augmented reality. This automated scene combination strategy provides users with a fast, efficient and accurate way to build complex three-dimensional scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] For better understanding and implementation, the technical solution of the present application is described in detail below with reference to the accompanying drawings.
[0017] Figure 1 A schematic diagram of a flow chart of a scene construction method based on three-dimensional reconstruction provided in Example 1 of the present application; Figure 2 A schematic diagram of a process for preprocessing a two-dimensional image using a scene construction method based on three-dimensional reconstruction provided in Example 1 of the present application; Figure 3 A schematic diagram of a process for automatically identifying and annotating different objects and regions in an image in a scene construction method based on three-dimensional reconstruction provided in Example 1 of the present application; Figure 4 A schematic diagram of a process for storing a reconstructed three-dimensional model as a sub-scene model in a scene building method based on three-dimensional reconstruction provided in Example 1 of the present application. DETAILED DESCRIPTION
[0018] In order to further explain the technical means and effects taken by the present invention to achieve the predetermined invention purpose, exemplary embodiments will be described in detail here, and examples thereof are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are only examples of methods and systems consistent with some aspects of the present application as detailed in the attached claims.
[0019] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms of "a", "said" and "the" used in this application and the appended claims are also intended to include plural forms unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.
[0020] The specific implementation methods, features and effects of the present invention are described in detail below in conjunction with the accompanying drawings and preferred embodiments.
[0021] Example 1 See also Figure 1-Figure 4 This embodiment provides a scene construction method and system based on 3D reconstruction, which effectively solves the technical bottleneck of real-time rendering of large-scale complex scenes and improves the efficiency and accuracy of scene construction through multi-view image acquisition, preprocessing, semantic segmentation and 3D reconstruction algorithm based on bundle adjustment.
[0022] The present invention provides a scene construction method based on three-dimensional reconstruction, comprising the following steps: S1, collect multi-view 2D images of real-world environments through drones, mobile phone cameras and image acquisition devices; Specifically, to obtain multi-perspective two-dimensional images of real-world environments, drones can be used to achieve bird's-eye view photography of large areas and obtain panoramic images; mobile phone cameras can be used for close-up, detailed photography. For the acquisition of multi-perspective images, algorithms can be used for image fusion and splicing. For example, by combining image registration and splicing algorithms in computer vision, images from multiple perspectives can be fused into a unified scene perspective for subsequent processing and analysis. The multi-perspective two-dimensional images obtained in this way can provide rich scene information and provide necessary input data for subsequent three-dimensional reconstruction and scene construction.
[0023] S2. Preprocess the acquired 2D images, such as denoising and color correction, to improve the quality of subsequent 3D reconstruction, and use semantic segmentation technology to automatically identify and annotate different objects and regions in the image, understand the semantic information in the image, and thus retain the category and functional characteristics of the object during the reconstruction process; Furthermore, the acquired two-dimensional image is preprocessed, specifically including: S11, removing noise from the two-dimensional image by bilateral filtering; Specifically expressed as: ; in, is the pixel value of the filtered image at point (x,y), is the original image at point The pixel value of is a spatial Gaussian function, is the spatial standard deviation, controlling the size of the filter, is the intensity Gaussian function, is the intensity standard deviation, controlling the similarity between pixel values, is the normalization factor; Specifically, bilateral filtering is a nonlinear filtering technique commonly used in image processing. It combines spatial Gaussian weights and intensity Gaussian weights to smooth images while retaining edge information. By cleverly combining spatial Gaussian functions and intensity Gaussian functions, it is possible to retain image edges while removing noise, making the processed image both natural and clear. The adjustable parameters of bilateral filtering enable the algorithm to adapt to different image characteristics and noise environments, and has good robustness. The effect of using bilateral filtering is to remove noise while maintaining the edge sharpness and details of the image, making the image more visually satisfactory, especially suitable for scenes where edge information needs to be retained.
[0024] S12, adjusting the gain of each channel through the grayscale world hypothesis algorithm to achieve color correction, and then adjusting the color according to the grayscale card in the scene or according to the information of the image itself through the white balance algorithm; The specific gray world hypothesis algorithm is expressed as: ;in is the average brightness of the entire image, is the average brightness of channel c; The specific white balance algorithm is expressed as: ; where I(c) is a channel of the original image, is the value of this channel after white balancing; Specifically, color correction using the grayscale world hypothesis algorithm and the white balance algorithm can significantly improve the color consistency of the image under different lighting conditions, thereby improving the realism and visual quality of the image; while the white balance algorithm further adjusts the color according to the grayscale card in the scene or the information of the image itself to ensure that the white objects in the image appear to be truly white. The advantage of using the grayscale world hypothesis and the white balance algorithm for color correction is that they can automatically adjust the image color to adapt to different lighting conditions, thereby ensuring the consistency and realism of the image color. By calculating and adjusting the RGB channel gain of the image to achieve color balance, the white object appears to be truly white in the image, which can improve the visual quality of the acquired image, improve the naturalness and accuracy of the color, and have wide applicability and strong robustness.
[0025] S13, registering the acquired two-dimensional image by using a feature matching algorithm and an image registration algorithm.
[0026] Specifically, a feature detection algorithm is used to detect key points and descriptors, and then matching point pairs are found by matching descriptors; a transformation matrix for aligning one image to another image is solved by a transformation matrix; the transformation matrix solution includes a homography matrix and an affine transformation; The homography matrix is expressed as: ; where x and is a pair of matching points, is a robust kernel function; The affine transformation is expressed as: ;in and are the coordinates of the matching point pairs.
[0027] Specifically, through feature point matching, accurate correspondence can be found between different images, making the registration result very accurate; different feature detection algorithms and transformation matrix solution methods, such as homography matrix and affine transformation, can also be selected according to different application requirements to adapt to different scenarios and requirements.
[0028] Furthermore, the use of semantic segmentation technology to automatically identify and label different objects and regions in the image specifically includes: S21, extract features from the preprocessed image through convolutional layers and pooling layers, and use a fully convolutional network for pixel-level classification; S22, use the softmax algorithm to classify each pixel, and enlarge the feature map to the original size through upsampling and deconvolution technology to generate a segmentation map; S23. Aggregate pixel labels into coherent object regions through threshold processing and connected component analysis, and identify and locate the boundaries of different objects.
[0029] Specifically, by extracting features through convolutional layers and pooling layers and then using a fully convolutional network for pixel-level classification, semantic segmentation technology can accurately identify objects and regions in the image; the softmax algorithm is used to classify each pixel to ensure the accuracy of classification, while the application of upsampling and deconvolution techniques allows the resolution of the segmentation result to match the original image, retaining the detailed information of the image; finally, through threshold processing and connected component analysis, the scattered pixel labels are aggregated into continuous regions, achieving clear recognition and positioning of object boundaries, thereby visually generating an accurate object segmentation map.
[0030] S3, according to the annotated two-dimensional image data, using a multi-viewpoint geometric three-dimensional reconstruction algorithm based on bundle adjustment, reconstructing a three-dimensional model from the two-dimensional image and storing the reconstructed three-dimensional model as a sub-scene model; Each sub-scene model contains its position, size and material information; Furthermore, the method of using a multi-viewpoint geometric 3D reconstruction algorithm based on bundle adjustment to reconstruct a 3D model from a 2D image and storing the reconstructed 3D model as a subscene model specifically includes: S31, matching corresponding feature points between images of different viewing angles, determining the positions of the feature points in different images, and determining the intrinsic and extrinsic parameters of the camera through a known calibration plate image; S32. Using multi-view feature point matching and camera parameters, the three-dimensional coordinates of each feature point in the space are determined by triangulation method; the triangulation method utilizes the different positions of the same feature point in different images when the feature point is observed from different viewpoints, thereby calculating its actual position in the three-dimensional space.
[0031] The triangulation formula is expressed as solving a linear system , where A is the matrix constructed based on the corresponding point coordinates and the camera projection matrix, and P is the homogeneous coordinates of the three-dimensional point.
[0032] S33, optimizing the three-dimensional coordinates and camera parameters by minimizing the reprojection error, and updating the parameter estimation value using the least square method, wherein the reprojection error refers to the difference between the three-dimensional point projected back to the two-dimensional image and the actual feature point position; The minimized reprojection error is expressed as: ; in, is the location of the feature point in the image, is the corresponding 3D point position, are the camera parameters, represents the projection function, is a robust kernel function used to handle outliers; S34, storing the optimized three-dimensional point coordinates and image camera parameters as a sub-scene model.
[0033] Specifically, by using a multi-viewpoint geometric 3D reconstruction algorithm based on bundle adjustment, 3D models can be accurately reconstructed from 2D images and stored as sub-scene models. First, the corresponding points of the image under different viewing angles are determined by feature point matching technology, and then the camera parameters are calibrated using the calibration plate image to ensure the accuracy of the reconstruction process. Then, the position of each feature point in 3D space is calculated by triangulation method, combined with multi-view feature point matching and accurate camera parameters. In addition, the accuracy of 3D point position and camera parameters is further improved by the optimization process of minimizing reprojection error, in which the application of robust kernel function effectively handles outliers and enhances the stability of the algorithm; finally, the optimized 3D point coordinates and camera parameters are used to construct sub-scene models, providing high-quality 3D data for subsequent graphics rendering, virtual reality, augmented reality and other applications. The advantages of the whole process are its high precision, high degree of automation, and the ability to effectively handle complex scenes, providing users with a reliable 3D reconstruction solution.
[0034] Furthermore, the use of the least squares method to update the parameter estimates specifically includes: ; where X is the design matrix, Y is the observation vector, is the parameter vector.
[0035] S4, automatically selecting a suitable stored sub-scenario model by specifying the required scenario type and specific requirements, and performing a preliminary combination according to the automatically selected instructions to form an initial scenario model; Furthermore, the automatically selecting a suitable stored sub-scene model and performing a preliminary combination according to the automatically selected instruction to form an initial scene model specifically includes: By defining keywords and tags, the stored sub-scenario models are indexed. According to the project requirement analysis results, intelligent retrieval is used to perform keyword matching and tag screening to find sub-scenario models that meet the requirements. Using computer graphics technology, the selected sub-scene models are spatially analyzed and calculated to determine the relative position and orientation of the sub-scene models in the virtual scene, and the models are automatically arranged and combined through a constraint-based layout algorithm.
[0036] Specifically, the advantages of the automated sub-scene model selection and combination process are its efficiency, accuracy and scalability. Through intelligent retrieval and computer graphics technology, manual operations can be greatly reduced and the speed and quality of scene construction can be improved. At the same time, the constraint-based layout algorithm can ensure the rationality and diversity of model combinations, while the real-time preview and fine-tuning functions provide an intuitive interactive method, allowing users to easily adjust the scene to achieve the desired effect.
[0037] S5, automatically making minor adjustments to the initial scene model to generate a target scene description file of the target scene model, wherein the target scene description file includes information for rendering and displaying the scene; Furthermore, the initial scene model is automatically fine-tuned by adding or removing objects, modifying materials, and adjusting lighting for training. Specifically, computer graphics techniques are used to automatically add or remove objects from a scene; Modify the material parameters by adjusting parameters such as reflectivity in the Blinn-Phong lighting model; adjust the material properties of the object surface, such as color, texture, roughness, etc., to match the style of the target scene; Use lighting models, a physically based rendering technique, to simulate real-world lighting effects and optimize lighting conditions in the scene, including adding or adjusting the position, intensity, and type of light sources.
[0038] In the physical rendering technology, the illumination calculation using the BRDF illumination model is specifically expressed as: ; in, is the Fresnel coefficient, which represents the reflectivity of the material. is the normal distribution function, which describes the effect of surface roughness on lighting. is a geometric function that describes the degree to which light is blocked by a surface. is the surface roughness, N is the surface normal, L is the direction vector from the surface point to the light source, V is the direction vector from the surface point to the camera, is the intensity of the light source.
[0039] Furthermore, after automatically making minor adjustments to the initial scene model and generating the target scene description file of the target scene model, the generated target scene model is optimized by combining the generative adversarial network and the variational autoencoder to make the 3D model more realistic, diversified, and with richer details and expressiveness. The specific contents include: The loss function of the variational autoencoder and the loss function of the generative adversarial network are combined to obtain a joint optimization goal, which is used to obtain a more realistic and more atmospheric 3D model. Specifically, the reconstruction ability of the variational autoencoder and the fidelity improvement ability of the generative adversarial network are utilized. The joint loss function is expressed as: ; in, and is the weight coefficient that balances the two loss terms, Expressed as reconstruction loss, it is used to measure the similarity between the 3D point cloud reconstructed by the latent variable z and the original point cloud; It is expressed as KL divergence, which is used to measure the distribution of latent variables. With the prior distribution The difference between Indicates that it encourages the discriminator to correctly identify the real 3D model. It is expressed as encouraging the generator to produce realistic fake 3D models. Through the loss function of the generative adversarial network, the generator learns to produce 3D models with rich details and high-quality surfaces.
[0040] Specifically, the Generative Adversarial Network (GAN) can learn the style and structure of the target object through an adversarial mechanism, thereby improving the visual expressiveness and realism of the generated 3D model. At the same time, combined with the variational autoencoder (VAE) technology, the model can be modeled and optimized in the latent space to achieve model generation and reconstruction. This technology can generate smoother and more continuous latent representations, improving the diversity and flexibility of the template scene model after generation.
[0041] Furthermore, a scene construction method based on three-dimensional reconstruction further includes: S6, according to the target scene description file for generating the target scene model, checking whether the generated target scene description file complies with preset rules, and deploying the model that complies with the rules to the virtual platform; Furthermore, a scene construction method based on three-dimensional reconstruction also includes: significantly improving the speed of scene construction through automated three-dimensional reconstruction technology and intelligent scene combination strategy; using real two-dimensional materials for three-dimensional reconstruction to make the generated scene more realistic.
[0042] The present embodiment provides a scene construction method based on 3D reconstruction, which aims to solve the challenges encountered by traditional technologies in real-time rendering of large-scale complex scenes. The method first collects multi-view 2D images of real-world environments through devices such as drones and mobile phone cameras. These images are preprocessed, including denoising and color correction, to improve the quality of subsequent 3D reconstruction. In the preprocessing step, bilateral filtering is used to smooth the image and retain edge information, while the grayscale world hypothesis and white balance algorithm are used to adjust the image color to ensure color consistency under different lighting conditions; then, the image is accurately registered through feature matching algorithms and image registration algorithms to ensure that multi-view images can be correctly aligned. With the help of semantic segmentation, objects and regions in the image are automatically identified and annotated to understand the semantic information in the image.
[0043] Subsequently, a multi-viewpoint geometric 3D reconstruction algorithm based on bundle adjustment is used to reconstruct 3D models from 2D images and store these models as sub-scene models. In this process, the 3D coordinates and camera parameters are optimized by minimizing the reprojection error to ensure the accuracy of the reconstructed 3D model.
[0044] In the sub-scene model selection and combination stage, the intelligent retrieval technology is used to automatically select the appropriate sub-scene model according to the project requirements, and the computer graphics technology is used for spatial analysis and calculation to automatically arrange and combine the models. Finally, the initial scene model is slightly adjusted, including adding or removing objects, modifying materials and adjusting lighting, to generate the target scene description file of the target scene model. In order to further improve the realism and diversity of the model, the generated target scene model is optimized by combining the generative adversarial network and the variational autoencoder to obtain a more realistic three-dimensional model, making the model more detailed and expressive. After verification and inspection to ensure that the generated target scene description file meets the preset rules, the model is deployed to the virtual platform, providing high-quality three-dimensional scenes for applications such as virtual reality and augmented reality. This automated three-dimensional reconstruction and scene building process not only improves efficiency, but also ensures the realism and visual quality of the scene.
[0045] Example 2 The present embodiment provides a scene building system based on three-dimensional reconstruction, which is applied to a scene building based on three-dimensional reconstruction, and is used to form an automated three-dimensional scene building system, which can efficiently reconstruct realistic three-dimensional scenes from two-dimensional images of the real world, and provide high-quality three-dimensional data for applications such as virtual reality and augmented reality.
[0046] A scene construction system based on three-dimensional reconstruction, specifically comprising: The image acquisition module uses drones, mobile phone cameras and image acquisition devices to obtain multi-view 2D images of real-world environments. It can achieve bird's-eye view photography of large areas and close-up detailed photography, providing rich scene information for 3D reconstruction. The image preprocessing module performs preprocessing operations such as denoising and color correction on the collected two-dimensional images, including bilateral filtering technology for noise removal, and grayscale world hypothesis and white balance algorithm for color correction; The feature extraction and image registration module uses a feature detection algorithm to detect key points and descriptors, and finds matching point pairs through a feature matching algorithm; Among them, the multi-view images are accurately aligned through the image registration algorithm of homography matrix and affine transformation; The semantic segmentation module uses the deep learning technology of the fully convolutional network to perform pixel-level semantic segmentation, automatically identify and annotate different objects and regions in the image, and understand the semantic information in the image; The 3D reconstruction and optimization module uses a multi-viewpoint geometric 3D reconstruction algorithm based on bundle adjustment to reconstruct a 3D model from a 2D image and stores the reconstructed 3D model as a subscene model.
[0047] The scene combination and adjustment module automatically selects appropriate sub-scene models through intelligent retrieval technology and performs preliminary combination to form an initial scene model. It is used to make minor adjustments to the initial scene model, including adding or removing objects, modifying materials and adjusting lighting, as well as generating target scene description files, and finally deploying the model that meets the rules to the virtual platform. This also includes using generative adversarial networks and variational autoencoders to optimize the initial scene model to improve the realism and diversity of the model.
[0048] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment as above, it is not used to limit the present invention. Any technical personnel in this field can make some changes or modify the technical contents disclosed above into equivalent embodiments without departing from the scope of the technical solution of the present invention. However, any brief modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.
Claims
1. A scene construction method based on three-dimensional reconstruction, characterized in that: The steps include: Collect multi-view 2D images of the environment through drones, mobile phone cameras, and image acquisition devices; Preprocess the acquired two-dimensional images and use semantic segmentation technology to automatically identify and annotate different objects and areas in the images and understand the semantic information in the images; According to the annotated two-dimensional image data, a three-dimensional model is reconstructed from the two-dimensional image using a multi-viewpoint geometric three-dimensional reconstruction algorithm based on bundle adjustment, and the reconstructed three-dimensional model is stored as a sub-scene model; The method utilizes a multi-viewpoint geometric 3D reconstruction algorithm based on bundle adjustment to reconstruct a 3D model from a 2D image and stores the reconstructed 3D model as a subscene model, specifically comprising: Match the corresponding feature points between images of different perspectives, determine the positions of the feature points in different images, and determine the intrinsic and extrinsic parameters of the camera through the known calibration plate image; Using multi-view feature point matching and camera parameters, the three-dimensional coordinates of each feature point in space are determined by triangulation method; Optimize the 3D coordinates and camera parameters by minimizing the reprojection error, which refers to the difference between the 3D point projected back to the 2D image and the actual feature point position, and use the least squares method to update the parameter estimates; The minimized reprojection error is expressed as: ; in, is the location of the feature point in the image, is the corresponding 3D point position, are the camera parameters, represents the projection function, is a robust kernel function used to handle outliers; According to the optimized coordinates of the 3D points, the camera parameters of the image are stored as a sub-scene model.
2. The scene construction method based on three-dimensional reconstruction according to claim 1, characterized in that: The acquired two-dimensional image is preprocessed, specifically including: The noise of the two-dimensional image is removed by bilateral filtering; the bilateral filtering is expressed as: ; in, is the pixel value of the filtered image at point (x,y), is the original image at point The pixel value of is a spatial Gaussian function, is the spatial standard deviation, controlling the size of the filter, is the intensity Gaussian function, is the intensity standard deviation, controlling the similarity between pixel values, is the normalization factor; The grayscale world hypothesis algorithm is used to adjust the gain of each channel for color correction, and then the white balance algorithm is used to adjust the color according to the grayscale card in the scene or according to the information of the image itself; The gray world hypothesis algorithm is expressed as: ;in is the average brightness of the entire image, is the average brightness of channel c; The white balance algorithm is expressed as: ; where I(c) is a channel of the original image, is the value of this channel after white balancing; The acquired two-dimensional images are registered by feature matching algorithm and image registration algorithm; Specifically, the method includes: using a feature detection algorithm to detect key points and descriptors, and then finding matching point pairs by matching descriptors; solving a transformation matrix for aligning one image to another image by using a transformation matrix; the transformation matrix solution includes a homography matrix and an affine transformation; The homography matrix is expressed as: ; where x and is a pair of matching points, is a robust kernel function; The affine transformation is expressed as: ;in and are the coordinates of the matching point pairs.
3. The scene construction method based on three-dimensional reconstruction according to claim 1, characterized in that: The use of semantic segmentation technology to automatically identify and label different objects and regions in the image specifically includes: Extract features from the preprocessed image through convolutional and pooling layers, and use a fully convolutional network for pixel-level classification; Use the softmax algorithm to classify each pixel, and use upsampling and deconvolution techniques to enlarge the feature map to its original size to generate a segmentation map; Through threshold processing and connected component analysis, pixel labels are aggregated into coherent object regions, and the boundaries of different objects are identified and located.
4. The scene construction method based on three-dimensional reconstruction according to claim 1, characterized in that: The method of using the least square method to update the parameter estimation value specifically includes: ; where X is the design matrix, Y is the observation vector, is the parameter vector.
5. The scene construction method based on three-dimensional reconstruction according to claim 1, characterized in that: Also includes: By specifying the required scene type and specific requirements, the appropriate stored sub-scene model is automatically selected and preliminarily combined according to the automatically selected instructions to form an initial scene model; Among them, the stored sub-scene models are indexed by defining keywords and tags, and according to the results of project demand analysis, intelligent retrieval is used to perform keyword matching and tag screening to find sub-scene models that meet the requirements; then computer graphics technology is used to perform spatial analysis and calculation on the selected sub-scene models to determine the relative position and orientation of the sub-scene models in the virtual scene, and the models are automatically arranged and combined through a constraint-based layout algorithm.
6. The scene construction method based on three-dimensional reconstruction according to claim 1, characterized in that: Also includes: The initial scene model is automatically fine-tuned to generate a target scene description file of the target scene model, wherein the target scene description file includes information for rendering and displaying the scene.
7. The scene construction method based on three-dimensional reconstruction according to claim 6, characterized in that: In the process of automatically making minor adjustments to the initial scene model, training is performed by adding or removing objects, modifying materials, and adjusting lighting. Specifically, computer graphics techniques are used to automatically add or remove objects from a scene; Adjust parameters in the Blinn-Phong lighting model to modify the material parameters; Use lighting models and physically based rendering techniques to simulate real-world lighting effects and optimize lighting conditions in the scene; In the physical rendering technology, the illumination calculation using the BRDF illumination model is specifically expressed as: ; in, is the Fresnel coefficient, which represents the reflectivity of the material. is the normal distribution function, which describes the effect of surface roughness on lighting. is a geometric function that describes the degree to which light is blocked by the surface. is the surface roughness, N is the surface normal, L is the direction vector from the surface point to the light source, V is the direction vector from the surface point to the camera, is the intensity of the light source.
8. The scene construction method based on three-dimensional reconstruction according to claim 6, characterized in that: After automatically making minor adjustments to the initial scene model and generating a target scene description file for the target scene model, the target scene model generated by optimizing the generated scene model by combining the generative adversarial network and the variational autoencoder includes: fusing the loss function of the variational autoencoder and the loss function of the generative adversarial network to obtain a joint optimization target for obtaining a three-dimensional model; The joint loss function is expressed as: ; in, and is the weight coefficient that balances the two loss terms, Expressed as reconstruction loss, it is used to measure the similarity between the 3D point cloud reconstructed by the latent variable z and the original point cloud; Expressed as KL divergence, it is used to measure the distribution of latent variables With prior distribution The difference between Indicates that it encourages the discriminator to correctly identify the real 3D model. This is expressed as encouraging the generator to produce realistic fake 3D models.
9. The scene construction method based on three-dimensional reconstruction according to claim 1, characterized in that: Also includes: According to the target scenario description file of the generated target scenario model, it is checked whether the generated target scenario description file complies with the preset rules, and the model that complies with the rules is deployed to the virtual platform.
10. A scene construction system based on three-dimensional reconstruction, applied to a scene construction method based on three-dimensional reconstruction as claimed in any one of claims 1 to 9, characterized in that: It includes image acquisition module, image preprocessing module, feature extraction and image registration module, semantic segmentation module, 3D reconstruction and optimization module and scene combination and adjustment module; The image acquisition module acquires multi-view two-dimensional images of the environment through drones, mobile phone cameras and image acquisition devices; The image preprocessing module performs preprocessing operations on the collected two-dimensional image, specifically including noise removal by bilateral filtering technology, and color correction by grayscale world hypothesis and white balance algorithm; The feature extraction and image registration module uses a feature detection algorithm to detect key points and descriptors, and finds matching point pairs through a feature matching algorithm; Among them, the multi-view images are accurately aligned through the image registration algorithm of homography matrix and affine transformation; The semantic segmentation module uses the deep learning technology of the fully convolutional network to perform pixel-level semantic segmentation, automatically identify and annotate different objects and regions in the image, and understand the semantic information in the image; The three-dimensional reconstruction and optimization module reconstructs a three-dimensional model from a two-dimensional image based on a multi-viewpoint geometric three-dimensional reconstruction algorithm of bundle adjustment, and stores the reconstructed three-dimensional model as a sub-scene model; The scene combination and adjustment module automatically selects appropriate sub-scene models through intelligent retrieval technology, performs preliminary combination to form an initial scene model, and then makes minor adjustments to the initial scene model, including adding or removing objects, modifying materials and adjusting lighting, generating a target scene description file, and then deploying the model that meets the rules to the virtual platform; This also includes optimizing the initial scene model using generative adversarial networks and variational autoencoders.
Citation Information
Patent Citations
Three-dimensional reconstruction method and system based on multi-view vision
CN117315138A
Satellite three-dimensional reconstruction method in complex illumination environment
CN119399344A
Scene three-dimensional reconstruction method and system, electronic equipment and storage medium
CN119722979A
Method, apparatus, and storage medium for three-dimensional reconstruction of buildings based on missing point cloud data
US20240257462A1
Selective inclusion of images in a three-dimensional model of a scene
WO2024199655A1
Cited By
Method and device for generating three-dimensional live-action model
CN120163927A
Method and device for generating three-dimensional real scene model
CN120163927B
Assembly part intelligent classification and attribute association method and device based on deep learning
CN120411662A
Visual three-dimensional reconstruction method and system of structure prior, equipment and medium
CN120510306A
Building design review system based on augmented reality technology
CN120747425A