Three-dimensional modeling method based on multi-view image fusion and AI semantic material decoupling
Through the 3D modeling method of multi-perspective image fusion and AI semantic material decoupling, the problems of texture seam elimination and insufficient material decoupling in the existing technology are solved, high-precision texture alignment and material restoration are achieved, and the application of models in virtual environments is supported.
Patent Information
- Application Number
- CN202510931923.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-10-17
AI Technical Summary
Existing 3D modeling technology relies on manual intervention to eliminate texture seams and color transitions during the image fusion process. In particular, texture blur and color disparity problems are prone to occur in complex material scenes. In addition, the ability to decouple object semantic information from material properties is insufficient, making it difficult to support refined material editing and light and shadow simulation.
A 3D modeling method based on multi-view image fusion and AI semantic material decoupling is adopted. Through an improved pyramid block matching algorithm, an adaptive overlapping area weighted fusion strategy, a dual-channel deep learning network and a generative adversarial network, combined with a cross-modal attention mechanism, the texture alignment accuracy is improved and the material properties are restored.
In the case of ancient building digitization, the texture alignment accuracy has been improved to 0.3 pixels, and the material property restoration error is less than 8%. It supports weathering simulation and virtual restoration of the model in the Unity/Unreal engine, improving the overall quality and consistency of the model.
Smart Images

Figure CN120807795A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of three-dimensional modeling, in particular to a three-dimensional modeling method based on multi-view image fusion and AI semantic material decoupling. BACKGROUND
[0002] As a core supporting technology in the fields of digital twinning, cultural heritage protection, and smart city, three-dimensional modeling technology has shown a trend of developing from traditional geometric modeling to intelligent and semantic modeling in recent years. Although existing three-dimensional reconstruction methods based on multi-view images achieve efficient real scene modeling, there are still two major bottlenecks: first, the elimination of texture seams and color transition in the image fusion process relies on manual intervention, especially in complex material scenes such as ancient building wooden structures and stone surfaces, which easily causes texture blurring and color difference faults; second, existing algorithms lack sufficient decoupling ability for object semantic information and material properties, making it difficult for the model to support fine material editing and light simulation.
[0003] In Chinese Patent CN120088409A, three-dimensional laser scanning modeling relies on point cloud data but cannot automatically distinguish tile and brick wall materials, and oblique photography modeling can capture the appearance but loses physical properties such as material reflectivity and roughness. Therefore, the present application proposes a three-dimensional modeling method based on multi-view image fusion and AI semantic material decoupling to solve the problems in the prior art. SUMMARY
[0004] To overcome the shortcomings of the prior art, the present application proposes a three-dimensional modeling method based on multi-view image fusion and AI semantic material decoupling. In the case of ancient building digitization, the texture alignment accuracy can be improved to 0.3 pixel level, the material property restoration error is less than 8%, and the model can directly perform weathering simulation and virtual repair operations in the Unity / Unreal engine, providing a new technical paradigm for the digital protection and active utilization of cultural heritage.
[0005] The technical solution of the present application is as follows: a three-dimensional modeling method based on multi-view image fusion and AI semantic material decoupling, comprising the following steps:
[0006] Step one: first, an improved pyramid block matching algorithm is used to adopt an adaptive overlapping region weighted fusion strategy in the image preprocessing stage;
[0007] Step two: dynamically adjust the pixel gray value by using the inverse distance weight function to eliminate the exposure difference and joint gap between multi-view images;
[0008] Step three: In the construction of a dual-channel deep learning network, the spatial geometry channel extracts building structure features through a graph convolution network, including the implementation of semantic segmentation of roof and beam components;
[0009] Step four: In the dual-channel deep learning network, the material decoupling channel uses an adversarial generative network to separate diffuse reflection, specular reflection, and normal map physical material properties from image high-frequency components;
[0010] Step five: Finally, by introducing a cross-modal attention mechanism, the geometric topological information and material features are mutually corrected in the latent space, and a hierarchical three-dimensional model carrying semantic labels and PBR material parameters is finally output.
[0011] Further improvement lies in that in the step one, the original image is subjected to multi-scale pyramid decomposition to generate image levels of different resolutions, the top layer of the pyramid is the lowest resolution image, and the bottom layer is the highest resolution image, each resolution level image is segmented into fixed-size blocks, feature extraction is performed on each block, matching results of different resolution levels are fused to obtain a globally optimal matching relationship, and geometric transformation correction is performed on the matching point pairs.
[0012] Further improvement lies in that in the step one, in the image preprocessing stage, the overlapping area between different images is determined according to the matching point pairs, the weight distribution of each image block in the overlapping area is calculated, the weight is calculated according to the relative position and quality of the image block in the overlapping area, the weight is smoothed by using a Gaussian kernel function, and the image blocks in the overlapping area are weighted and averaged according to the calculated weight to generate a fused image.
[0013] Further improvement lies in that in the step two, first, all multi-view images are accurately registered, the corresponding points between the images are aligned, exposure correction is performed on each image, and the overlapping area of each pair of images is determined.
[0014] Further improvement lies in that in the step two, by the weighted average process, the brightness difference between different images is smoothly transitioned, thereby eliminating the stitching gap, and a seamless panoramic image is generated by fusing the images.
[0015] Further improvement lies in that in the step three, multi-view image data is collected and preprocessed, including cropping, scaling, and normalization, and the structure features of buildings, including roofs and beams, are labeled for training a deep learning model, a dual-channel deep learning network is designed, one channel processes spatial geometry information, and the other channel processes semantic information, the spatial geometry channel uses a graph convolution network, and the semantic channel uses a convolutional neural network.
[0016] Further improvement lies in that in the step three, the image is extracted features by using a semantic channel adopting a convolutional neural network, the semantic channel adopting the convolutional neural network can capture local features and global features in the image, and the semantic segmentation of the building component is realized by the semantic channel adopting the convolutional neural network.
[0017] Further improvement lies in that in the step four, an adversarial generative network is designed, including a generator and a discriminator, the generator is responsible for generating physical material properties such as diffuse reflection, specular reflection and normal map, the discriminator is responsible for distinguishing the generated material properties and the real material properties, the generator learns to generate high-quality physical material properties, and a loss function is defined, including an adversarial loss and a perceptual loss.
[0018] Further improvement lies in that in the step four, the adversarial generative network is trained using multi-view image data, and the network parameters are optimized through continuous iteration until the generated material properties meet the predetermined quality standard.
[0019] Further improvement lies in that in the step five, the cross-modal attention mechanism is used to fuse the geometric features and the material features, the geometric features and the material features are corrected in the fusion process, the hierarchical three-dimensional model is constructed according to the fused feature information, the hierarchical model represents different levels of details from rough geometric shapes to fine material textures, and the three-dimensional modeling is completed in this way.
[0020] Compared with the prior art, the method has the following advantages: the texture alignment accuracy in the digital case of the ancient building can be improved to 0.3 pixel level, the material attribute restoration error is less than 8%, and the model can be directly weathered and simulated in the Unity / Unreal engine and virtual repair operation is supported, a new technical paradigm is provided for the digital protection and active utilization of cultural heritage, semantic segmentation is performed by using deep learning, semantic labels can be assigned to different parts of the model, the cross-modal attention mechanism is introduced, the geometric topological information and the material features are corrected in the latent space, and the overall quality and consistency of the model are improved. BRIEF DESCRIPTION OF DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.
[0022] Figure 1 The step flow chart of the present application. DETAILED DESCRIPTION
[0023] With reference to the drawings of the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described, obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative work belong to the protection scope of the present application.
[0024] In the description of the present application, it should be noted that, unless otherwise explicitly specified and limited, the terms installation, connection, and connection should be understood in a broad sense, for example, it can be fixed connection, or detachable connection, or integrally connected; it can be mechanical connection, or electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, or the communication inside two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.
[0025] In the file CN120088409A, the modeling efficiency and reliability in complex scenes are significantly improved, the potential conflict of the multi-modal input data is intelligently analyzed through a dynamic arbitration mechanism, the preset priority rules or physical rules are automatically applied for content arbitration, the delay and error of manual intervention decision are effectively avoided; when the uploaded image and text description have attribute contradictions, the arbitration result is generated according to the quantization indexes such as data clarity and user instruction explicitness, so that the fusion accuracy of cross-modal data is improved; the joint modeling of semantics and space is realized by integrating a large language model and a neural radiation field, so as to ensure that the generated three-dimensional model meets the double standards in geometric precision and semantic consistency, and the real-time embedding of the physical simulation engine can identify structural defects in the initial modeling stage, but in three-dimensional laser scanning modeling, it depends on point cloud data and cannot automatically distinguish tile and brick wall materials, so in the present application, the multi-view image fusion can capture the all-around details of the target object and scene, thereby generating a high-precision three-dimensional model, and the AI semantic material decoupling technology can separate high-quality physical basic rendering material parameters including diffuse reflection, specular reflection and normal map from the image, so that the model has a realistic visual effect.
[0026] Referring to Figure 1 The embodiment of the present application discloses a three-dimensional modeling method based on multi-view image fusion and AI semantic material decoupling, comprising the following steps:
[0027] Step one: first, an improved pyramid block matching algorithm is used to adopt an adaptive overlapping region weighted fusion strategy in the image preprocessing stage;
[0028] Step two: the pixel gray value is dynamically adjusted by using an inverse distance weight function to eliminate the exposure difference and splicing gap between multi-view images;
[0029] Step three: In the construction of the dual-channel deep learning network, the spatial geometry channel extracts the building structure features through the graph convolution network, including the implementation of semantic segmentation of roof and beam components;
[0030] Step four: In the dual-channel deep learning network, the material decoupling channel uses the generative adversarial network to separate the diffuse reflection, specular reflection and normal map physical material attributes from the high-frequency components of the image;
[0031] Step five: Finally, by introducing the cross-modal attention mechanism, the geometric topological information and material features are mutually corrected in the latent space, and finally the hierarchical three-dimensional model carrying semantic labels and PBR material parameters is output.
[0032] In step one, the original image is subjected to multi-scale pyramid decomposition to generate image levels of different resolutions, the top layer of the pyramid is the lowest resolution image, and the bottom layer is the highest resolution image. Each resolution level image is divided into fixed size blocks, and feature extraction is performed on each block, including SIFT, SURF and ORB features. In the current resolution level, the image blocks of different angles are matched, and the improved matching algorithm is used to quickly find the best matching pair. Pyramid down-sampling matching is used to pass the matching results from the current resolution level to the next lower resolution level. The matching results are refined on the low resolution level to improve the robustness and accuracy of the matching. The matching results are fused to obtain the globally optimal matching relationship. The matching point pairs are geometrically transformed and corrected.
[0033] In step one, in the image preprocessing stage, the overlapping area between different images is determined according to the matching point pairs, the weight distribution of each image block in the overlapping area is calculated, the weight is calculated according to the relative position and quality of the image block in the overlapping area, the weight is smoothed using a Gaussian kernel function, the image blocks in the overlapping area are weighted and averaged according to the calculated weight, and the fused image is generated, wherein the boundary effect of the overlapping area is minimized. The fused image is quality evaluated to check for artifacts and discontinuities, and if necessary, the fusion strategy is adjusted and the fusion process is re-performed.
[0034] In step two, first, all multi-view images are accurately registered, aligning corresponding points between images, and performing exposure correction on each image to reduce differences in brightness and contrast between different images. This is achieved through histogram equalization and adaptive histogram equalization. The overlapping area of each pair of images is determined, obtained through feature matching and geometric transformation. In the overlapping area, a weight value is calculated for each pixel, which is inversely proportional to the distance of the pixel to the image boundary. The farther the distance, the lower the weight. The closer the distance, the higher the weight. For each pixel in the overlapping area, its gray value is dynamically adjusted according to its weight value. The gray value of this pixel is replaced by the weighted average of the gray values of its corresponding pixels in different images.
[0035] In step two, through the weighted average process, the brightness difference between different images is smoothly transitioned, eliminating the stitching gap and generating a seamless panoramic image. All images after weight adjustment are fused to generate a seamless panoramic image. Quality assessment is performed on the fused image to check for any remaining exposure differences and stitching marks, and the weight function is further optimized.
[0036] In step three, multi-view image data is collected and preprocessed, including cropping, scaling, and normalization. The structural features of buildings, including roofs and columns, are labeled for training the deep learning model. A dual-channel deep learning network is designed, with one channel processing spatial geometric information and the other channel processing semantic information. The spatial geometric channel uses a graph convolution network, and the semantic channel uses a convolutional neural network. The graph convolution network is used to model the spatial geometric information of buildings, capturing the relationship between nodes and building edges, and extracting the structural features of buildings, including the geometric shapes and positions of roofs and columns.
[0037] In step three, the semantic channel uses a convolutional neural network to extract features from the image. The semantic channel uses a convolutional neural network to capture local and global features in the image. Through the semantic channel using a convolutional neural network, semantic segmentation of building components is achieved, distinguishing different building parts such as roofs and columns. The spatial geometric features extracted by the graph convolution network and the semantic features extracted by the semantic channel using a convolutional neural network are fused using feature fusion techniques, including attention mechanisms and gating mechanisms, to integrate the two features and improve the accuracy of segmentation. The dual-channel deep learning network is trained using the labeled dataset, and the cross-entropy loss function is used to optimize the loss function for semantic segmentation.
[0038] In step four, an adversarial generative network is designed, including a generator and a discriminator. The generator is responsible for generating physical material properties such as diffuse reflection, specular reflection, and normal map. The discriminator is responsible for distinguishing between generated material properties and real material properties. The generator generates preliminary physical material properties from image high-frequency components. Through iterative training, the performance of the generator is optimized so that the generated material properties are increasingly close to the true values. In the process of adversarial training, the generator and the discriminator compete with each other. The generator tries to generate more and more realistic material properties, while the discriminator tries to distinguish between true and false material properties. In this way, the generator learns to generate high-quality physical material properties. A loss function is defined, which typically includes adversarial loss and perceptual loss. Additional loss terms, including L1 and L2 loss, can be introduced to improve the stability of generated materials.
[0039] In step four, the adversarial generative network is trained using multi-view image data. Through continuous iteration, network parameters are optimized until the generated material properties meet the predetermined quality standards.
[0040] In step five, cross-modal attention mechanism is used to fuse geometric features and material features. In the fusion process, geometric features and material features correct each other. Based on the fused feature information, a hierarchical three-dimensional model is constructed. The hierarchical model represents different levels of details, from rough geometry to fine material texture. The constructed hierarchical three-dimensional model is optimized, and the final hierarchical three-dimensional model is output in a standard format. The model is applied to various fields to complete three-dimensional modeling.
[0041] This three-dimensional modeling method based on multi-view image fusion and AI semantic material decoupling captures multi-view images using multiple cameras to capture target objects and scenes from different angles. Preprocessing operations such as denoising, color correction, and exposure adjustment are performed on the captured images. Feature extraction is performed, including SIFT, SURF, and ORB features. An improved pyramid block matching algorithm is used for feature matching. Through an adaptive overlapping region weighting fusion strategy, the images are integrated to generate three-dimensional point cloud or mesh models. Deep learning models, including convolutional neural networks, are applied to perform semantic segmentation on the images. Adversarial generative networks are used to separate physical material properties such as diffuse reflection, specular reflection, and normal map from image high-frequency components. Based on the fused point cloud or mesh data, three-dimensional reconstruction is performed to generate an initial three-dimensional model. The model is smoothed, holes are filled, and topology is optimized to improve the quality of the model. Post-processing such as lighting, shading, and reflection is performed on the three-dimensional model to enhance its realism. High-resolution texture maps are applied to the model to further enhance the visual effect. Cross-modal attention mechanism is introduced to allow geometric topology information and material features to correct each other in latent space. Finally, a hierarchical three-dimensional model is constructed to represent different levels of details, and semantic labels are assigned to different parts of the model to generate PBR material parameters.
[0042] The above description is only the preferred embodiment of the present application, and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A 3D modeling method based on multi-view image fusion and AI semantic material decoupling, characterized in that: The following steps are involved: Step 1: First, an adaptive overlapping area weighted fusion strategy is adopted in the image preprocessing stage through the improved pyramid block matching algorithm; Step 2: Dynamically adjust the pixel grayscale value through the inverse distance weight function to eliminate the exposure differences and stitching gaps between multi-view images; Step 3: When building a dual-channel deep learning network, the spatial geometry channel extracts building structural features through a graph convolutional network, including semantic segmentation of eaves and beams and columns. Step 4: In the dual-channel deep learning network, the material decoupling channel uses a generative adversarial network to separate the physical material properties of diffuse reflection, specular reflection, and normal map from the high-frequency components of the image; Step 5: Finally, by introducing a cross-modal attention mechanism, the geometric topology information and material features are mutually corrected in the latent space, and finally a hierarchical 3D model with semantic labels and PBR material parameters is output.
2. The 3D modeling method based on multi-view image fusion and AI semantic material decoupling according to claim 1, characterized in that: In step 1, the original image is subjected to multi-scale pyramid decomposition to generate image levels of different resolutions, with the top layer of the pyramid being the lowest resolution image and the bottom layer being the highest resolution image. The image of each resolution level is divided into blocks of fixed size, and features are extracted for each block. The matching results of different resolution levels are fused to obtain the global optimal matching relationship, and geometric transformation correction is performed on the matching point pairs.
3. The 3D modeling method based on multi-view image fusion and AI semantic material decoupling according to claim 1, characterized in that: In step one, after the correction is completed, in the image preprocessing stage, the overlapping area between different images is determined based on the matching point pairs, the weight distribution of each image block in the overlapping area is calculated, the weight is calculated based on the relative position and quality of the image block in the overlapping area, the weight is smoothed using a Gaussian kernel function, and the image blocks in the overlapping area are weighted averaged according to the calculated weights to generate a fused image.
4. The 3D modeling method based on multi-view image fusion and AI semantic material decoupling according to claim 1, characterized in that: In the second step, all multi-view images of the fused image are accurately registered, corresponding points between the images are aligned, exposure correction is performed on each image, and the overlapping area of each pair of images is determined.
5. The 3D modeling method based on multi-view image fusion and AI semantic material decoupling according to claim 1, characterized in that: In the second step, the brightness difference between different images is smoothly transitioned through the weighted averaging process, thereby eliminating the stitching gaps and generating a fused image. All the weighted images are fused to generate a seamless panoramic image.
6. The 3D modeling method based on multi-view image fusion and AI semantic material decoupling according to claim 1, characterized in that: In the step three, multi-view image data of the panoramic image is collected and preprocessed, including cropping, scaling and normalization, and the structural features of the building, including eaves and beams, are marked for training the deep learning model. A dual-channel deep learning network is designed, in which one channel processes spatial geometric information and the other channel processes semantic information. The spatial geometry channel uses a graph convolutional network and the semantic channel uses a convolutional neural network.
7. The 3D modeling method based on multi-view image fusion and AI semantic material decoupling according to claim 1, characterized in that: In the step three, a convolutional neural network is used to extract features from the image using a semantic channel. The convolutional neural network used in the semantic channel can capture local and global features in the image, and semantic segmentation of building components is achieved through the convolutional neural network used in the semantic channel.
8. The 3D modeling method based on multi-view image fusion and AI semantic material decoupling according to claim 1, characterized in that: In the step four, a generative adversarial network is designed, which includes a generator and a discriminator. The generator is responsible for generating physical material properties such as diffuse reflection, specular reflection and normal map, and the discriminator is responsible for distinguishing the generated material properties from the real material properties. The generator learns to generate high-quality physical material properties and defines loss functions, including adversarial loss and perceptual loss.
9. The 3D modeling method based on multi-view image fusion and AI semantic material decoupling according to claim 1, characterized in that: In step 4, the adversarial generative network is trained using multi-view image data, and the network parameters are optimized through continuous iteration until the generated material properties meet the predetermined quality standards.
10. The 3D modeling method based on multi-view image fusion and AI semantic material decoupling according to claim 1, characterized in that: In step five, a cross-modal attention mechanism is finally used to fuse the geometric features and material features. During the fusion process, the geometric features and material features are mutually corrected, and a hierarchical three-dimensional model is constructed based on the fused feature information. The hierarchical model represents different levels of details, from rough geometric shapes to fine material textures, thereby completing three-dimensional modeling.
Citation Information
Patent Citations
Three-dimensional automatic modeling method and system based on AI large model technology
CN120088409A
Cited By
Three-dimensional structure modeling method based on zero sample semantic segmentation and physical attribute estimation
CN121414992A
Three-dimensional structure modeling method based on zero-shot semantic segmentation and physical property estimation
CN121414992B