A Multimodal Method for Generating 3D Models

Through multimodal data fusion technology, deep learning is used to generate three-dimensional models, which solves the problems of high cost, long cycle and low accuracy in traditional methods, and realizes the generation of high-precision and strong sense of reality, supporting real estate decoration and sales.

CN118864757BActive Publication Date: 2025-07-22MOZHI VISION (HUZHOU) TECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410895103.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-05
Publication Date
2025-07-22
Estimated Expiration
2044-07-05

AI Technical Summary

Technical Problem

Traditional three-dimensional modeling methods of real estate rely on a single data source, are costly, long cycles and complex operations, making it difficult to accurately display three-dimensional structures, textures and details. The existing technology cannot provide comprehensive real estate information.

Method used

By obtaining various modal data of the property, including visible light images, point cloud data and CAD floor drawings, deep learning technology is used to extract three-dimensional structures, plan layouts and visual features, and feature fusion is used for PointNet++, ViT and Transformer models to generate a three-dimensional model.

Benefits of technology

It improves the accuracy and reality of the three-dimensional model, provides detailed data support for real estate decoration design and sales, and realizes efficient and automated three-dimensional model generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118864757B_ABST
    Figure CN118864757B_ABST
Patent Text Reader

Abstract

This application relates to the field of 3D modeling technology. Specifically, it discloses a method for generating a 3D model in a multi-modal manner. By obtaining various modal data of a property, including visible light images, point cloud data, and CAD floor plan drawings, and using deep learning technology to extract the 3D structural information, floor plan layout information, and visual features of the property from the various modal data, a 3D model of the property is intelligently generated based on the attention fusion features of the multi-modal data of the property. In this way, the advantages of various data sources can be fully utilized to improve the accuracy and realism of the 3D model, providing detailed data support for links such as the decoration design and sales of the property.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of 3D modeling technology, and more specifically, to a method for generating 3D models with multiple modalities. Background Art

[0002] With the rapid development and wide application of information technology, unprecedented changes have taken place in all walks of life. In the real estate field, the application of information technology is becoming increasingly widespread, especially in the display and modeling of real estate, where the demand is growing vigorously.

[0003] Traditional real estate display methods, such as flat drawings and static pictures, although they can meet the basic needs of home buyers and developers to a certain extent, have obvious limitations. Specifically, although flat drawings can present the floor plan of a property, they cannot intuitively display the 3D structure and spatial layout of the property, making it difficult for designers to accurately grasp the sense of space and proportion during the design process. Static pictures, on the other hand, often can only show a partial or specific perspective of the property, lacking integrity and dynamics, and unable to provide home buyers with comprehensive real estate information.

[0004] In the existing technology, the 3D modeling of real estate mainly relies on professional 3D scanning equipment or manual modeling software, which has disadvantages such as high cost, long cycle, and complex operation. In addition, these 3D modeling methods often only use a single data source for modeling, such as CAD house type drawings, which can provide basic spatial information but are lacking in texture, color, and details, limiting the accuracy of the model. Therefore, a method for generating 3D models with multiple modalities is expected. Summary of the Invention

[0005] To solve the above technical problems, this application is proposed. The embodiments of this application provide a method for generating 3D models with multiple modalities, which obtains various modality data of a property, including visible light images, point cloud data, and CAD house type drawings, and uses deep learning technology to extract the 3D structure information, floor plan information, and visual features of the property from the various modality data, thereby intelligently generating a 3D model of the property based on the attention fusion features of the multi-modal data of the property. In this way, the advantages of various data sources can be fully utilized to improve the accuracy and realism of the 3D model, providing detailed data support for links such as the decoration design and sales of the property.

[0006] Correspondingly, according to one aspect of this application, a method for generating 3D models with multiple modalities is provided, which includes:

[0007] Obtain 3D point cloud data of a property, a CAD house type drawing of the property, and multiple visible light images of the property;

[0008] Based on the three-dimensional point cloud data of the property, extract three-dimensional features of the property to obtain a three-dimensional spatial structure feature vector of the property;

[0009] Extract the planar layout features of the CAD floor plan of the property to obtain a planar layout feature vector of the property;

[0010] Perform image semantic feature extraction and feature fusion on the multiple visible-light images of the property to obtain a global visual feature vector of the property;

[0011] Generate a three-dimensional model of the property based on the multi-modal fusion features of the three-dimensional spatial structure feature vector, the planar layout feature vector, and the global visual feature vector of the property.

[0012] In the above multi-modal three-dimensional model generation method, based on the three-dimensional point cloud data of the property, extracting three-dimensional features of the property to obtain a three-dimensional spatial structure feature vector of the property includes: passing the three-dimensional point cloud data of the property through a three-dimensional feature extractor of the property based on the PointNet++ model to obtain the three-dimensional spatial structure feature vector of the property.

[0013] In the above multi-modal three-dimensional model generation method, extracting the planar layout features of the CAD floor plan of the property to obtain a planar layout feature vector of the property includes: passing the CAD floor plan of the property through a floor plan feature extractor of the property based on a convolutional neural network model to obtain the planar layout feature vector of the property.

[0014] In the above multi-modal three-dimensional model generation method, performing image semantic feature extraction and feature fusion on the multiple visible-light images of the property to obtain a global visual feature vector of the property includes: passing the multiple visible-light images of the property through a visual feature extractor of the property based on the ViT model respectively to obtain multiple visual feature vectors of the property, and fusing the multiple visual feature vectors to obtain the global visual feature vector of the property.

[0015] In the above multi-modal three-dimensional model generation method, generating a three-dimensional model of the property based on the multi-modal fusion features of the three-dimensional spatial structure feature vector, the planar layout feature vector, and the global visual feature vector of the property includes: performing attention fusion on the three-dimensional spatial structure feature vector, the planar layout feature vector, and the global visual feature vector of the property to obtain a multi-modal fusion feature vector of the property; generating the three-dimensional model of the property based on the multi-modal fusion feature vector of the property.

[0016] In the above multi-modal three-dimensional model generation method, attention fusion is performed on the three-dimensional spatial structure feature vector of the property, the plane layout feature vector of the property, and the global visual feature vector of the property to obtain a multi-modal fusion feature vector of the property, including: passing the three-dimensional spatial structure feature vector of the property, the plane layout feature vector of the property, and the global visual feature vector of the property through a multi-modal feature fusion device based on a Transformer model to perform feature fusion to obtain the multi-modal fusion feature vector of the property.

[0017] In the above multi-modal three-dimensional model generation method, based on the multi-modal fusion feature vector of the property, a three-dimensional model of the property is generated, including: inputting the multi-modal fusion feature vector of the property into a three-dimensional model generator based on an adversarial generation network to obtain the three-dimensional model of the property.

[0018] In the above multi-modal three-dimensional model generation method, based on the multi-modal fusion feature vector of the property, a three-dimensional model of the property is generated, including: enhancing the regression characteristics of the multi-modal fusion feature vector of the property based on auxiliary analysis and description to obtain an optimized multi-modal fusion feature vector of the property; inputting the optimized multi-modal fusion feature vector of the property into a three-dimensional model generator based on an adversarial generation network to obtain the three-dimensional model of the property.

[0019] In the above multi-modal three-dimensional model generation method, enhancing the regression characteristics of the multi-modal fusion feature vector of the property based on auxiliary analysis and description to obtain an optimized multi-modal fusion feature vector of the property includes: multiplying the multi-modal fusion feature vector of the property by the generation weight matrix of the adversarial generation network to obtain an intermediate feature vector; concatenating the multi-modal fusion feature vector of the property and the intermediate feature vector to obtain a concatenated feature vector; multiplying the concatenated feature vector by a first weight matrix and then adding a first bias vector to obtain an encoded concatenated feature vector; activating the encoded concatenated feature vector through a sigmoid function to obtain an activation value; multiplying the intermediate feature vector by the result of subtracting the activation value from 1 element-wise to obtain a weighted intermediate feature vector; multiplying the activation value by the multi-modal fusion feature vector of the property element-wise to obtain a weighted multi-modal fusion feature vector of the property; subtracting the multi-modal fusion feature vector of the property from the intermediate feature vector element-wise to obtain a difference feature vector; adding the weighted intermediate feature vector and the weighted multi-modal fusion feature vector of the property and then dividing by the difference feature vector to obtain a second intermediate feature vector; passing the second intermediate feature vector through a ReLU function to obtain the optimized multi-modal fusion feature vector of the property.

[0020] Compared with the prior art, the multi-modal three-dimensional model generation method provided by this application obtains various modal data of a property, including visible light images, point cloud data, and CAD floor plan drawings, and uses deep learning technology to extract the three-dimensional structural information, floor plan layout information, and visual features of the property from the various modal data, so as to intelligently generate a three-dimensional model of the property based on the attention fusion features of the multi-modal data of the property. In this way, the advantages of various data sources can be fully utilized to improve the accuracy and realism of the three-dimensional model, providing detailed data support for links such as the decoration design and sales of the property. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] By describing the embodiments of this application in more detail with reference to the accompanying drawings, the above and other objects, features, and advantages of this application will become more apparent. The accompanying drawings are used to provide a further understanding of the embodiments of this application and constitute a part of the specification, and are used together with the embodiments of this application to explain this application, and do not constitute a limitation to this application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.

[0022] Figure 1 It is a flowchart of the multi-modal three-dimensional model generation method according to the embodiment of this application.

[0023] Figure 2 It is a schematic diagram of the architecture of the multi-modal three-dimensional model generation method according to the embodiment of this application.

[0024] Figure 3 It is a flowchart of generating a three-dimensional model of a property based on the multi-modal fusion features of the three-dimensional space structure feature vector, the floor plan layout feature vector, and the global visual feature vector of the property in the multi-modal three-dimensional model generation method according to the embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] Next, the embodiments of this application will be described in more detail with reference to the accompanying drawings, and the above and other objects, features, and advantages of this application will become more apparent. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments of this application. It should be understood that this application is not limited by the example embodiments described here. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this invention.

[0026] As described in the above background art, in the current technical field, the three-dimensional modeling of real estate mainly relies on professional three-dimensional scanning equipment and complex manual modeling software. However, these methods generally face challenges such as high costs, long cycles, and cumbersome operations. In addition, these three-dimensional modeling methods mainly rely on a single data source, such as CAD house type drawings. Although such methods can show the spatial structure to a certain extent, there are obvious deficiencies in the presentation of textures, colors, and details, thus limiting the comprehensiveness and accuracy of the generated three-dimensional model and making it difficult for the final result to achieve the expected effect. To address the above technical problems, the technical concept of this application is to obtain various modal data of real estate, including visible light images, point cloud data, and CAD house type drawings, and use deep learning technology to extract the three-dimensional structure information, floor plan information, and visual features of real estate from the various modal data, so as to intelligently generate a three-dimensional model of real estate based on the attention fusion features of multi-modal real estate data. In this way, the advantages of various data sources can be fully utilized to improve the accuracy and realism of the three-dimensional model and provide detailed data support for links such as the decoration design and sales of real estate.

[0027] Figure 1 It is a flowchart of a multi-modal three-dimensional model generation method according to an embodiment of the present application. Figure 2 It is a schematic diagram of the architecture of a multi-modal three-dimensional model generation method according to an embodiment of the present application. As Figure 1 and Figure 2 shown, the multi-modal three-dimensional model generation method according to an embodiment of the present application includes the steps of: S110, obtaining real estate three-dimensional point cloud data, real estate CAD house type drawings, and multiple real estate visible light images; S120, extracting three-dimensional features of real estate based on the real estate three-dimensional point cloud data to obtain a real estate three-dimensional spatial structure feature vector; S130, extracting the floor plan features of the real estate CAD house type drawing to obtain a real estate floor plan feature vector; S140, performing image semantic feature extraction and feature fusion on the multiple real estate visible light images to obtain a real estate global visual feature vector; S150, generating a real estate three-dimensional model based on the multi-modal fusion features of the real estate three-dimensional spatial structure feature vector, the real estate floor plan feature vector, and the real estate global visual feature vector.

[0028] In the above multi-modal three-dimensional model generation method, in step S110, three-dimensional point cloud data of a property, a CAD floor plan of the property, and multiple visible light images of the property are obtained. It should be understood that point cloud data is a set of points in three-dimensional space, directly recording the three-dimensional coordinate information of the object's surface. For a property, the three-dimensional point cloud data can capture detailed information such as the spatial structure, size, and shape of the property, providing direct spatial data support for constructing a three-dimensional model. The CAD floor plan is a plan view of the property, usually containing detailed information such as the layout of rooms, dimensions, and the positions of doors and windows, which plays an important role in understanding the planar layout and internal spatial relationships of the property. And multiple visible light images of the property can capture visual information such as the appearance, texture, and color of the property, providing rich visual feature support for constructing a three-dimensional model. In the technical solution of this application, by comprehensively using these data of different modalities to capture the spatial structure, planar layout, and visual features of the property, the actual situation of the property can be more comprehensively reflected, which helps to improve the accuracy and realism of three-dimensional model construction.

[0029] In the above multi-modal three-dimensional model generation method, in step S120, based on the three-dimensional point cloud data of the property, three-dimensional features of the property are extracted to obtain a three-dimensional spatial structure feature vector of the property. In a specific example of this application, the encoding method for extracting three-dimensional features of the property based on the three-dimensional point cloud data to obtain a three-dimensional spatial structure feature vector is to pass the three-dimensional point cloud data of the property through a three-dimensional feature extractor of the property based on the PointNet++ model to obtain the three-dimensional spatial structure feature vector. It should be understood that the three-dimensional spatial structure information of the property is the basis for constructing its three-dimensional model. However, due to the disorderliness of point cloud data, that is, the points in the three-dimensional coordinate point set have no specific order, this poses a challenge to traditional deep learning models. The PointNet++ model is a deep learning network designed for point cloud data, which can effectively process disorderly point cloud data and capture the three-dimensional spatial structure information of the property through multi-level point cloud feature extraction and abstraction. Specifically, the PointNet++ model adopts a hierarchical sampling and grouping strategy. First, the point cloud data is sampled to obtain point cloud subsets of different scales; then, each point cloud subset is grouped to form local neighborhoods; next, a multi-layer perceptron (MLP) is used to extract features from each local neighborhood; finally, the local features of different scales are fused to obtain the three-dimensional spatial structure feature of the property. In this way, the PointNet++ model can accurately capture the three-dimensional spatial structure information of the property, including the shape, size, and positional relationship of rooms, providing strong support for constructing a high-precision three-dimensional model.

[0030] In the above multi-modal three-dimensional model generation method, in step S130, the planar layout features of the real estate CAD floor plan are extracted to obtain a real estate planar layout feature vector. In a specific example of the present application, the encoding method for extracting the planar layout features of the real estate CAD floor plan to obtain a real estate planar layout feature vector is to pass the real estate CAD floor plan through a floor plan feature extractor based on a convolutional neural network model to obtain the real estate planar layout feature vector. It should be understood that the planar layout information of a property determines the distribution, size, and mutual relationship of the rooms in the three-dimensional model of the property. As the main source of planar layout information, the rich details contained in the real estate CAD floor plan play a key role in understanding the internal structure of the property. Therefore, in order to accurately extract the planar layout features of the real estate CAD floor plan, in the technical solution of the present application, a convolutional neural network model with excellent performance in the field of image processing and computer vision is used to process the real estate CAD floor plan. Specifically, the floor plan feature extractor based on the convolutional neural network model performs sliding convolution operations on the real estate CAD floor plan to extract the planar layout features of the property, including planar layout information such as the shape and size of the rooms, the relative positional relationship between the rooms, and the positions of doors and windows, providing an important reference basis for constructing an accurate three-dimensional model.

[0031] In the above multi-modal three-dimensional model generation method, in step S140, image semantic feature extraction and feature fusion are performed on the multiple visible-light images of the property to obtain a global visual feature vector of the property. In a specific example of the present application, the encoding method for performing image semantic feature extraction and feature fusion on the multiple visible-light images of the property to obtain a global visual feature vector of the property is to respectively pass the multiple visible-light images of the property through a property visual feature extractor based on the ViT model to obtain multiple property visual feature vectors, and fuse the multiple property visual feature vectors to obtain the global visual feature vector of the property. It should be understood that the visual features of the property play an important role in constructing a three-dimensional model with a high degree of realism and rich details. The visible-light images of the property are an important visual representation of the exterior and interior environments of the property, containing rich visual information such as color, texture, and lighting. In order to fully extract the visual features of the property, in the technical solution of the present application, the ViT (Vision Transformer) model, which performs well in the field of machine vision, is used as the property visual feature extractor to process the multiple visible-light images of the property. Specifically, the ViT model processes image data by introducing the Transformer structure, and directly applies the Encoder in the Transformer model to the image feature extraction part, realizing cross-domain migration from natural language processing to the visual field. Compared with traditional CNN models, the ViT model has stronger global feature extraction capabilities and can better model long-range dependencies in images. Here, the property visual feature extractor based on the ViT model performs global semantic feature extraction on each visible-light image of the property through the self-attention mechanism to capture visual features such as the appearance, texture, and color of the property. Furthermore, the visual features of multiple visible-light images of the property are fused by means of feature concatenation to obtain a global visual feature representation of the property. In this way, rich visual information is introduced into the construction of the three-dimensional model of the property, enabling the constructed three-dimensional model to maintain a high degree of realism and rich details in appearance.

[0032] In the above multi-modal three-dimensional model generation method, in step S150, a three-dimensional model of the property is generated based on the multi-modal fusion features of the three-dimensional spatial structure feature vector, the planar layout feature vector, and the global visual feature vector of the property. Among them, Figure 3 It is a flowchart for generating a three-dimensional model of a property based on the multi-modal fusion features of the three-dimensional spatial structure feature vector, the planar layout feature vector, and the global visual feature vector of the property in the multi-modal three-dimensional model generation method according to an embodiment of the present application. As Figure 3As shown, the step S150 includes: S151, performing attention fusion on the three-dimensional spatial structure feature vector of the property, the planar layout feature vector of the property, and the global visual feature vector of the property to obtain a multi-modal fusion feature vector of the property; S152, generating the three-dimensional model of the property based on the multi-modal fusion feature vector of the property.

[0033] Specifically, in step S151, attention fusion is performed on the three-dimensional spatial structure feature vector of the property, the planar layout feature vector of the property, and the global visual feature vector of the property to obtain a multi-modal fusion feature vector of the property. In a specific example of the present application, the encoding method for performing attention fusion on the three-dimensional spatial structure feature vector of the property, the planar layout feature vector of the property, and the global visual feature vector of the property to obtain a multi-modal fusion feature vector of the property is to perform feature fusion on the three-dimensional spatial structure feature vector of the property, the planar layout feature vector of the property, and the global visual feature vector of the property through a multi-modal feature fusion device based on the Transformer model to obtain the multi-modal fusion feature vector of the property. It should be understood that during the generation process of the three-dimensional model of the property, the three-dimensional spatial structure feature vector of the property mainly describes the three-dimensional structure of the property, the planar layout feature vector of the property focuses on the planar layout of the rooms, and the global visual feature vector of the property captures the visual information of the appearance and internal environment of the property. By performing feature fusion on the three, the complementarity of multi-modal data can be fully utilized, thereby improving the accuracy and authenticity of the three-dimensional model generation. In the technical solution of the present application, in order to achieve effective multi-modal feature fusion, a multi-modal feature fusion device based on the Transformer model is used to perform attention fusion processing on the three. Specifically, the Transformer model has achieved remarkable results in many fields with its powerful self-attention mechanism and global modeling ability. Here, the multi-head self-attention mechanism of the Transformer model is used to capture the dependency relationships between the three-dimensional spatial structure feature vector of the property, the planar layout feature vector of the property, and the global visual feature vector of the property, and perform feature weighted fusion accordingly, effectively integrating and correlating the feature vectors from different modalities, so as to obtain a more comprehensive and accurate multi-modal fusion feature vector of the property.

[0034] Specifically, the step S152 generates the real estate three-dimensional model based on the real estate multimodal fusion feature vector. In a specific example of the present application, the implementation method of generating the real estate three-dimensional model based on the real estate multimodal fusion feature vector is to input the real estate multimodal fusion feature vector into a three-dimensional model generator based on a generative adversarial network to obtain the real estate three-dimensional model. It should be understood that the generative adversarial network (GAN, Generative Adversarial Networks) has made breakthrough progress in the field of image generation with its unique generation and adversarial mechanism. In the technical solution of the present application, a generative adversarial network is used as a three-dimensional model generator, and the real estate multimodal fusion feature vector is used as input, and feature decoding is performed through a series of network layers to obtain the generation of the real estate three-dimensional model. Specifically, the generative adversarial network consists of two parts: a generator network and a discriminator network. The generator network is responsible for converting the input real estate multimodal fusion feature vector into a three-dimensional model, while the discriminator network is responsible for judging whether the generated three-dimensional model is real, that is, whether it is close to the appearance and internal structure of the real estate. Through the continuous game and optimization between the generator and the discriminator, the accuracy and authenticity of the three-dimensional model generation can be gradually improved. During the training process, a large amount of real estate data is first used to pre-train the GAN so that it can learn the appearance and internal structure characteristics of the real estate. Then, in the testing phase, the extracted multimodal fusion feature vector of the real estate is input into the trained GAN generator to quickly generate a three-dimensional model that is highly similar to the real estate. In this way, not only can the complementarity of multimodal data be fully utilized to improve the accuracy and authenticity of the three-dimensional model generation, but also an efficient and automated three-dimensional model generation process can be achieved, which is of great significance to the digital and intelligent development of the real estate industry, and helps to improve user experience, promote transaction efficiency, and reduce costs.

[0035] In particular, in the technical solution of the present application, considering that the overall feature distribution manifold monotonicity of the real estate multimodal fusion feature vector is poor, this will make it difficult for the three-dimensional model generator to effectively learn the true distribution of the data, thereby affecting the accuracy of the generated results. That is, due to the poor overall feature distribution manifold monotonicity of the real estate multimodal fusion feature vector, the three-dimensional model generator may not be able to fully utilize the feature information of the data when generating the three-dimensional model, resulting in insufficient generation regression constraints of the three-dimensional model generator, thereby affecting the accuracy of the generated results, making it difficult for the three-dimensional model generator to accurately generate a three-dimensional model of the real estate. Therefore, in order to improve the accuracy of the generated results, in the technical solution of the present application, the regression characteristics of the real estate multimodal fusion feature vector are further enhanced based on the auxiliary analysis description to obtain an optimized real estate multimodal fusion feature vector.

[0036] Based on this, in a preferred embodiment of the present application, generating the three-dimensional model of the property based on the multi-modal fusion feature vector of the property includes: enhancing the regression characteristics of the multi-modal fusion feature vector of the property based on auxiliary analysis description to obtain an optimized multi-modal fusion feature vector of the property; inputting the optimized multi-modal fusion feature vector of the property into a three-dimensional model generator based on a generative adversarial network to obtain the three-dimensional model of the property.

[0037] Specifically, enhancing the regression characteristics of the multi-modal fusion feature vector of the property based on auxiliary analysis description to obtain an optimized multi-modal fusion feature vector of the property includes: multiplying the multi-modal fusion feature vector of the property by the generation weight matrix of the generative adversarial network to obtain an intermediate feature vector; splicing the multi-modal fusion feature vector of the property and the intermediate feature vector to obtain a spliced feature vector; multiplying the first weight matrix by the spliced feature vector and then adding the first bias vector to obtain an encoded spliced feature vector; activating the encoded spliced feature vector through a sigmoid function to obtain an activation value; multiplying the activation value subtracted from one by the intermediate feature vector element-wise to obtain a weighted intermediate feature vector; multiplying the activation value by the multi-modal fusion feature vector of the property element-wise to obtain a weighted multi-modal fusion feature vector of the property; subtracting the intermediate feature vector from the multi-modal fusion feature vector of the property element-wise to obtain a difference feature vector; adding the weighted intermediate feature vector and the weighted multi-modal fusion feature vector of the property and then dividing by the difference feature vector to obtain a second intermediate feature vector; passing the second intermediate feature vector through a ReLU function to obtain the optimized multi-modal fusion feature vector of the property.

[0038] That is, enhancing the regression characteristics of the multi-modal fusion feature vector of the property based on auxiliary analysis description to obtain an optimized multi-modal fusion feature vector of the property includes: enhancing the regression characteristics of the multi-modal fusion feature vector of the property based on auxiliary analysis description with the following optimization formula to obtain the optimized multi-modal fusion feature vector of the property, where the optimization formula is:

[0039]

[0040] where, v c represents the multi-modal fusion feature vector of the property, M is the generation weight matrix of the generative adversarial network, concat represents the splicing function, W1 represents the first weight matrix, b1 represents the first bias vector, represents matrix multiplication, sigmoid represents the S-shaped activation function, τ represents the activation value, ReLU represents the rectified linear unit function, and V′ represents the optimized multi-modal fusion feature vector of the property.

[0041] In the technical solution of this application, the manifold monotonicity of the overall feature distribution of the real estate multi-modal fusion feature vector is poor, resulting in insufficient generation regression constraint with respect to the overall feature distribution of the 3D model generator based on the adversarial generative network during generation regression, which affects the accuracy of the generation result.

[0042] Therefore, in the technical solution of this application, the regression characteristic enhancement of the real estate multi-modal fusion feature vector based on auxiliary analysis description is carried out. It performs auxiliary analysis and description of the geometric characteristics of the feature manifold of the real estate multi-modal fusion feature vector with the generation weight matrix of the adversarial generative network, and uses the difference in the feature manifold between the auxiliary description analysis result and the original real estate multi-modal fusion feature vector as a differential factor to support the description of the improved real estate multi-modal fusion feature vector for the predetermined generation result of the 3D model generator based on the adversarial generative network. Finally, an activation mechanism is added for activation to maintain the reinforcement of the distribution dependence with positive description. In this way, the manifold monotonicity of the overall feature layout of the real estate multi-modal fusion feature vector can be significantly improved, thereby enhancing its generation regression constraint with respect to the overall feature layout of the 3D model generator based on the adversarial generative network and improving the accuracy of the generation result.

[0043] In summary, the multi-modal generation 3D model method according to the embodiments of this application is elucidated. It obtains various modal data of the real estate, including visible light images, point cloud data, and CAD floor plan drawings, and uses deep learning technology to extract the 3D structure information, floor plan layout information, and visual features of the real estate from the various modal data, so as to intelligently generate the 3D model of the real estate based on the attention fusion features of the real estate multi-modal data. In this way, the advantages of various data sources can be fully utilized to improve the accuracy and realism of the 3D model, providing detailed data support for links such as the decoration design and sales of the real estate.

[0044] The basic principles of the present invention have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, advantages, effects, etc. mentioned in the present invention are only examples and not limitations, and it cannot be considered that these advantages, advantages, effects, etc. are essential for each embodiment of the present invention. In addition, the specific details of the above embodiments are only for the purposes of illustration and easy understanding, rather than limitations, and these details do not limit the present invention to necessarily adopt the above specific details to implement.

[0045] In the above embodiments, the descriptions of the various embodiments each have their own focuses. For parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments. In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the module division is only a logical function division, and there may be other division methods in actual implementation. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0046] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0047] In addition, in each of the embodiments of the present invention, the functional modules can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a combination of hardware and software functional modules.

[0048] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to cover all changes falling within the meaning and scope of the equivalent elements of the claims in the present invention. Any associated drawing marks in the claims should not be regarded as limiting the claimed rights.

[0049] In addition, it is obvious that the word "including" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units stated in the system claims can also be implemented by one unit through software or hardware.

[0050] Finally, it should be noted that the above description has been given for purposes of illustration and description. In addition, the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A multi-modal method for generating a three-dimensional model, characterized in that, Including: Obtaining three-dimensional point cloud data of a property, a CAD floor plan of the property, and multiple visible-light images of the property; Extracting three-dimensional features of the property based on the three-dimensional point cloud data of the property to obtain a three-dimensional spatial structure feature vector of the property; Extracting three-dimensional features of the property based on the three-dimensional point cloud data of the property to obtain a three-dimensional spatial structure feature vector of the property, including: Passing the three-dimensional point cloud data of the property through a three-dimensional feature extractor of the property based on the PointNet++ model to obtain the three-dimensional spatial structure feature vector of the property; Extracting the planar layout features of the CAD floor plan of the property to obtain a planar layout feature vector of the property; Performing image semantic feature extraction and feature fusion on the multiple visible-light images of the property to obtain a global visual feature vector of the property; Generating a three-dimensional model of the property based on the multi-modal fusion features of the three-dimensional spatial structure feature vector, the planar layout feature vector, and the global visual feature vector of the property; Among them, generating a three-dimensional model of the property based on the multi-modal fusion features of the three-dimensional spatial structure feature vector, the planar layout feature vector, and the global visual feature vector of the property, including: Passing the three-dimensional spatial structure feature vector, the planar layout feature vector, and the global visual feature vector of the property through a multi-modal feature fusion device based on the Transformer model to perform feature fusion to obtain a multi-modal fusion feature vector of the property; Inputting the multi-modal fusion feature vector of the property into a three-dimensional model generator based on an adversarial generative network to obtain the three-dimensional model of the property.

2. The multimodal three-dimensional model generation method according to claim 1, wherein Extracting the planar layout features of the CAD floor plan of the property to obtain a planar layout feature vector of the property, including: Passing the CAD floor plan of the property through a floor plan feature extractor based on a convolutional neural network model to obtain the planar layout feature vector of the property.

3. The multi-modal three-dimensional model generation method according to claim 2, wherein Performing image semantic feature extraction and feature fusion on the multiple visible-light images of the property to obtain a global visual feature vector of the property, including: Passing the multiple visible-light images of the property through a visual feature extractor of the property based on the ViT model respectively to obtain multiple visual feature vectors of the property, and fusing the multiple visual feature vectors to obtain the global visual feature vector of the property.

4. The multimodal three-dimensional model generation method according to claim 3, wherein Generating the three-dimensional model of the property based on the multi-modal fusion feature vector of the property, including: Performing regression characteristic enhancement based on auxiliary analysis description on the multi-modal fusion feature vector of the property to obtain an optimized multi-modal fusion feature vector of the property; Inputting the optimized multi-modal fusion feature vector of the property into a three-dimensional model generator based on an adversarial generative network to obtain the three-dimensional model of the property.

5. The multimodal three-dimensional model generation method according to claim 4, characterized in that, Performing regression characteristic enhancement based on auxiliary analysis description on the multi-modal fusion feature vector of the property to obtain an optimized multi-modal fusion feature vector of the property, including: Multiplying the multi-modal fusion feature vector of the property by the generation weight matrix of the adversarial generative network to obtain an intermediate feature vector; Concatenating the multi-modal fusion feature vector of the property and the intermediate feature vector to obtain a concatenated feature vector; Multiplying the first weight matrix by the concatenated feature vector and then adding the first bias vector to obtain an encoded concatenated feature vector; Activate the encoded splicing feature vector through the sigmoid function to obtain an activation value; Subtract the activation value from one and then multiply it element-wise with the intermediate feature vector to obtain a weighted intermediate feature vector; Multiply the activation value element-wise with the real estate multi-modal fusion feature vector to obtain a weighted real estate multi-modal fusion feature vector; Subtract the intermediate feature vector from the real estate multi-modal fusion feature vector element-wise to obtain a difference feature vector; Add the weighted intermediate feature vector and the weighted real estate multi-modal fusion feature vector, and then divide the result by the difference feature vector to obtain a second intermediate feature vector; Apply the ReLU function to the second intermediate feature vector to obtain the optimized real estate multi-modal fusion feature vector.