Three-dimensional scene automatic generation method and system based on scene semantics and geometric constraints
Through the cross-modal KanMiDiffusion denoising network model, combined with DeBERTa and Dinov2 visual encoder, the three-dimensional scene generation of the MiDiffusion model is optimized, solving the problem of insufficient three-dimensional spatial structure modeling in the existing technology, and achieving higher semantic and geometric accuracy and diversity.
Patent Information
- Application Number
- CN202510756822.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-09
AI Technical Summary
The existing MiDiffusion model lacks explicit modeling of three-dimensional spatial structures, and the interaction mechanism between discrete semantics and continuous geometry is not perfect, resulting in low semantic and geometric accuracy of scene generation.
The cross-modal KanMiDiffusion denoising network model is adopted, combined with the DeBERTa text encoder, Dinov2 visual encoder and the Kolmogolov-Arnold network, a three-dimensional scene is generated through the fusion of geometry and text, and the Transformer decoder is used to optimize the interaction of semantics and geometric features.
Improve the semantic and geometric accuracy of scene generation, enhance the quality and diversity of scene generation, and demonstrate strong performance and generalization capabilities on different data sets.
Smart Images

Figure CN120278911A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of neural networks, and specifically relates to a method and system for automatically generating 3D scenes based on scene semantics and geometric constraints. Background Art
[0002] The 3D scene synthesis technology under realistic conditions has important research value in the fields of computer vision, robotics, and virtual reality. By constructing a high-fidelity virtual environment, this technology can not only provide training data for complex traffic scenes for autonomous driving systems, but also provide a diverse indoor environment test platform for the development of service robots. The MiDiffusion model has pioneered a new paradigm of hybrid discrete-continuous diffusion models. By introducing a two-dimensional floor plan and a structured object layout representation, the MiDiffusion model innovatively realizes the diffusion process in a hybrid domain that spans discrete semantics (such as furniture categories) and continuous geometry (such as position coordinates). However, the existing technologies have the following problems: First, the existing MiDiffusion model only relies on a two-dimensional plane as a spatial prior and lacks explicit modeling of the three-dimensional spatial structure.
[0003] Second, the interaction mechanism between discrete semantics and continuous geometry in the existing MiDiffusion model is not perfect enough, and the ability to understand semantics and extract geometric structures is insufficient.
[0004] The above problems lead to low semantic and geometric accuracy in scene generation. Summary of the Invention
[0005] The purpose of the embodiments of this application is to provide a method and system for automatically generating 3D scenes based on scene semantics and geometric constraints, which can solve the problem of low semantic and geometric accuracy in scene generation in the existing technologies.
[0006] To solve the above technical problems, this application is implemented as follows: In a first aspect, the embodiments of this application provide a method for automatically generating 3D scenes based on scene semantics and geometric constraints. The method includes: Build a preset cross-modal KanMiDiffusion denoising network model, and perform text encoding processing on latent variables according to the DeBERTa text encoder inside the KanMiDiffusion denoising network model to obtain text encoding features. The cross-modal KanMiDiffusion denoising network model includes: a first Kolmogorov-Arnold network, a second Kolmogorov-Arnold network, a third Kolmogorov-Arnold network, a Transformer decoder, a DeBERTa text encoder, and a Dinov2 visual encoder; Encode the geometric structure of the preset indoor objects according to the first Kolmogorov - Arnold network to obtain geometric encoding features; Add the geometric encoding features and the text encoding features to generate fused geometric and text features; Perform visual encoding on the preset indoor floor plan according to the Dinov2 visual encoder and perform feature splicing with the preset position encoding features to obtain fused visual and position features; Process the fused geometric and text features and the fused visual and position features according to the Transformer decoder to obtain the output of the Transformer decoder; Process the output of the Transformer decoder according to the second Kolmogorov - Arnold network and the third Kolmogorov - Arnold network to generate the classification distribution of semantic labels and the Gaussian mean of geometric features; Generate a 3D scene based on the classification distribution of semantic labels and the Gaussian mean of geometric features.
[0007] As an optional implementation manner of the first aspect of this application, the Kolmogorov - Arnold network includes multiple KAN Layer layers, and the mathematical expression is: ; Among them, represents the th layer of the KAN network, represents the th layer of the KAN network, represents the first layer of the KAN network, represents the 0th layer of the KAN network, represents the input vector, represents the input vector corresponding output result; Among them, each KAN network has an input of dimensions and an output of dimensions, represents the multiplication process, consists of the learnable activation function and the mathematical expression is: ; The mathematical expression of the Kolmogorov - Arnold network from the th layer to the th layer is: ; Among them, represents the output vector of the th layer, Indicates the output vector of the layer, and represents the matrix of the
[0008] As an alternative implementation of the first aspect of the present application, the DeBERTa text encoder includes: multiple layers of Transformer encoder layers connected in parallel, a word embedding layer, and a word vector representation layer. Each layer of the Transformer encoder layer includes multiple Transformer encoders connected in series. Among them, the word embedding layer is used to embed each position information, and the sentence embedding and word embedding are used as input vectors and respectively input into the multiple layers of Transformer encoder layers connected in parallel for encoding processing to generate the word vectors corresponding to each Transformer encoder layer.
[0009] As an alternative implementation of the first aspect of the present application, the Dinov2 visual encoder is based on the ViT architecture. The image is divided into multiple small patches to obtain an image patch sequence to capture the long-range dependencies in the image. The Dinov2 visual encoder includes multiple layers of Transformer modules to generate high-level feature representations rich in semantics.
[0010] As an alternative implementation of the first aspect of the present application, the Transformer decoder includes multiple Transformer blocks. Each Transformer block is sequentially connected to a first adaptive normalization layer, a multi-head self-attention layer, a second adaptive normalization layer, a multi-head cross-attention layer, and a feed-forward neural network layer. The output of the Transformer decoder is obtained by connecting the outputs of each self-attention head in the multi-head self-attention layer and performing a linear projection calculation. The mathematical expression of the output of the Transformer decoder is: ; where represents the input vector corresponding to the output of the Transformer decoder, are learnable parameters, represents the feature concatenation process, represents the weighted sum of N value vectors.
[0011] In a second aspect, the embodiments of the present application provide a three-dimensional scene automatic generation system based on scene semantics and geometric constraints. The system includes: A DeBERTa text encoder module for performing text encoding processing on latent variables to obtain text encoding features; The first Kolmogorov - Arnold module is used to encode the geometric structure of a preset indoor object to obtain geometric encoding features; The fusion feature generation module is used to add the geometric encoding features and the text encoding features to generate the geometric - text fusion features; The Dinov2 visual encoder module is used to perform visual encoding on a preset indoor floor plan and then splice the features with preset position encoding features to obtain visual - position fusion features; The Transformer encoder module is used to process the geometric - text fusion features and the visual - position fusion features to obtain the output of the Transformer decoder; The semantic label and geometric feature generation module is used to process the output of the Transformer decoder according to the second Kolmogorov - Arnold network and the third Kolmogorov - Arnold network to generate the classification distribution of semantic labels and the Gaussian mean of geometric features; The 3D scene generation module is used to generate a 3D scene based on the classification distribution of semantic labels and the Gaussian mean of geometric features.
[0012] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method in the first aspect are implemented.
[0013] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by the processor, the steps of the method in the first aspect are implemented.
[0014] Compared with the prior art, the beneficial effects of the 3D scene automatic generation method based on scene semantics and geometric constraints provided by the present invention are as follows: By integrating the DeBERTa text encoder and the Dinov2 visual encoder, and introducing the Kolmogorov - Arnold (KAN) network to optimize the geometric feature mapping, the cross - modal KanMiDiffusion algorithm is proposed, and the effectiveness of the KanMiDiffusion algorithm proposed by the present invention in improving the quality, efficiency, and diversity of scene generation is verified. It can improve the semantic and geometric accuracy of scene generation, and demonstrates the strong performance and generalization ability of the algorithm on different datasets. Description of the Drawings
[0015] Figure 1 is a flowchart of the 3D scene automatic generation method based on scene semantics and geometric constraints provided by the first embodiment of the present application; Figure 2It is a diagram of the Transformer model provided by the first embodiment of the present application; Figure 3 It is a diagram of the cross-modal KanMiDiffusion denoising network model provided by the first embodiment of the present application; Figure 4 It is a diagram of the ViT model provided by the first embodiment of the present application; Figure 5 It is a diagram of the DeBERTa text encoder model provided by the first embodiment of the present application; Figure 6 It is a diagram of the multi-modal feature interaction Transformer provided by the first embodiment of the present application; Figure 7 It is a diagram of the multi-head attention structure provided by the first embodiment of the present application; Figure 8 It is a diagram of a three-dimensional indoor scene one synthesized by using the cross-modal KanMiDiffusion algorithm provided by the first embodiment of the present application; Figure 9 It is a diagram of a three-dimensional indoor scene two synthesized by using the cross-modal KanMiDiffusion algorithm provided by the first embodiment of the present application; Figure 10 It is a diagram of a three-dimensional indoor scene three synthesized by using the cross-modal KanMiDiffusion algorithm provided by the first embodiment of the present application; Figure 11 It is a diagram of a three-dimensional indoor scene four synthesized by using the cross-modal KanMiDiffusion algorithm provided by the first embodiment of the present application; Figure 12 It is a diagram of a three-dimensional indoor scene five synthesized by using the cross-modal KanMiDiffusion algorithm provided by the first embodiment of the present application; Figure 13 It is an internal structure diagram of the three-dimensional scene automatic generation system based on scene semantics and geometric constraints provided by the second embodiment of the present application. Detailed implementation manners
[0016] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0017] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here. In addition, the "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally represents that the objects associated with each other are in an "or" relationship.
[0018] The following is a detailed description of the method for automatically generating a three-dimensional scene based on scene semantics and geometric constraints provided by the embodiment of the present application through specific embodiments and their application scenarios in conjunction with the accompanying drawings.
[0019] Example 1 See also Figure 1 , which is a flowchart of the method for automatically generating a three-dimensional scene based on scene semantics and geometric constraints provided by the present invention, includes steps S1 to S7.
[0020] Step S1: Build a preset cross-modal KanMiDiffusion denoising network model, and perform text encoding processing on the latent variables according to the DeBERTa text encoder inside the KanMiDiffusion denoising network model to obtain text encoding features, wherein the cross-modal KanMiDiffusion denoising network model includes: a first Kolmogorov-Arnold network, a second Kolmogorov-Arnold network, a third Kolmogorov-Arnold network, a Transformer decoder, a DeBERTa text encoder, and a Dinov2 visual encoder. Figure 5 Represents the DeBERTa text encoder model diagram provided by the present invention.
[0021] Specifically, the present invention first introduces the traditional MiDiffusion model, which pioneered a new paradigm of hybrid discrete-continuous diffusion model. The model innovatively implements the hybrid domain diffusion process across discrete semantics (such as furniture categories) and continuous geometry (such as location coordinates) by introducing two-dimensional floor plans and structured object layout representations. MiDiffusion is a hybrid discrete-continuous three-dimensional indoor scene synthesis diffusion model that assumes that each scene All in a central world frame, consisting of a two-dimensional plane and at most objects. Using a conditional vector to represent the floor plan. Each object in the scene has a semantic label , center of mass position , the bounding box size , the rotation angle around the vertical axis , expressed as: ; Among them, represents the planar rotation group, represents the rotation angle, ranging from 0 to 360 degrees, represents the matrix transpose.
[0022] All geometric features in the scene can be mapped to . Represented using a 3D indoor scene as , represents the given conditional input.
[0023] Set the last semantic label to an empty label to handle scenes containing less than objects, and define all-zero geometric properties for a "null" object.
[0024] Although object semantics and geometric properties are in different domains, the present invention defines a destruction process in hybrid diffusion, injecting specific-domain noise into and respectively. The mathematical expression for the noise is: ; Among them, represents the mixed probability distribution of semantic labels and geometric features, represents the probability distribution of obtaining the next-step semantic label given the previous-step semantic label for an object at step t, represents the probability distribution of obtaining the next-step semantic label given the previous-step semantic label for an object at step t.
[0025] Sampling is performed independently during training , , and the mathematical expression for the posterior distribution is: ; Among them, represents the object semantic label, represents the denoising process of geometric features, indicating the probability distribution of inferring the original feature and from the given noisy feature , represents the mixed denoising process.
[0026] During the backward diffusion process, design a network to calculate the probability distribution, which is consistent with the forward diffusion process. The mathematical expression is: ; Among them, represents that the neural network is used to fit the denoising process to obtain a specific probability distribution, and is used to fit the probability distribution, ( and represents the and joint conditional terms deduced backward from and . Among them, represents the given conditional input. For example, given a specific conditional input, a specific original feature can be deduced from it, such as the input process of a text image.
[0027] Furthermore, please refer to Figure 3 , which represents the cross-modal KanMiDiffusion denoising network model diagram provided by the present invention. This model diagram describes the denoising process, including latent variables in step and the predicted classification distribution , as well as the Gaussian mean . Sampling is performed through these distributions to obtain the latent variable in step .
[0028] In summary, the present invention uses the text encoding DeBERTa to encode the object category information. Secondly, the Kolmogorov-Arnold Network (KAN) is used to transfer geometric attributes to encode the object features, which are combined with the trainable semantic embeddings after bert encoding. Then, for each scene, the Dinov2 visual encoder is used as a planar feature extractor because the Dinov2 visual encoder can more easily capture the floor semantic information and better capture the floor placement area and floor boundary information. Finally, the planar features are added with position embeddings to form a conditional vector.
[0029] Step S2: According to the first Kolmogorov-Arnold network, encode the geometric structure of the preset indoor object to obtain geometric encoding features.
[0030] Specifically, the Multilayer Perceptron (MLP) is a fully connected feedforward neural network, which is the basic construction of a deep learning model and is usually used to approximate non-linear functions in machine learning. An MLP composed of layers can be described as a transformation matrix and an activation function The function, the mathematical expression is: ; Among them, represents the linear weight function of the th layer, represents the linear weight function of the th layer, represents the linear weight function of the 1st layer, represents the input vector, represents the output result of the multi-layer perceptron, represents the multiplication process.
[0031] The Kolmogorov-Arnold Network (KAN) provided by the present invention can replace the traditional multi-layer perceptron. Different from the traditional multi-layer perceptron MLP based on the generalized approximation theorem, KAN is a function approximation method inspired by the Kolmogorov-Arnold representation theorem. KAN has a fully connected structure similar to that of MLP, but different from MLP, MLP depends on the fixed activation function of each node, and KAN introduces a learnable activation function on the edge, fundamentally changing the neural network structure. Using the learnable one-dimensional spline function, the linear weight matrix is eliminated. Similar to MLP, a layer KAN can be described as the nesting of multiple KAN Layer layers, and the mathematical expression is: ; Among them, represents the th layer of the KAN network, represents the th layer of the KAN network, represents the 1st layer of the KAN network, represents the 0th layer of the KAN network, represents the input vector, represents the input vector corresponding output result.
[0032] Among them, each KAN network has dimensional input and dimensional output, represents the multiplication process, is composed of the learnable activation function , and the mathematical expression is: ; The mathematical expression of the Kolmogorov-Arnold network from the th layer to the th layer is: ; Among them, represents the output vector of the th layer, represents the output vector of the th layer, represents the th layer matrix of the KAN network.
[0033] The cross-modal KanMiDiffusion provided by the present invention uses KAN to transfer geometric attributes to encode object features and combines them with trainable semantic embeddings. At the same time, the output of the multi-modal feature interaction module Tranformer decoder is provided to two KAN networks to generate the classification distribution of semantic labels and the Gaussian mean of geometric features.
[0034] Step S3: Add the geometrically encoded features and the textually encoded features to generate fused geometric and text features.
[0035] Step S4: After visually encoding the preset indoor floor plan by the Dinov2 visual encoder, perform feature splicing with the preset position encoding features to obtain fused visual and position features.
[0036] Specifically, the Dinov2 visual encoder is based on the Vision Transformer (ViT) architecture, divides the image into multiple small patches to obtain an image patch sequence, and captures long-range dependencies in the image. The Dinov2 visual encoder includes multiple layers of Transformer modules to generate high-level feature representations rich in semantics.
[0037] Among them, the ViT model introduces the Transformer architecture into the field of computer vision, and especially shows its unique advantages when dealing with tasks such as image recognition. ViT divides the image into small patches, treats these patches as elements in the sequence, and uses the self-attention mechanism to construct the global feature representation of the image. This process is finally completed by a multi-layer perceptron.
[0038] Please refer to Figure 4 , which represents the ViT model diagram provided by the present invention. The structural design of the ViT model diagram gives it multiple advantages compared to traditional CNNs, as follows: 1) ViT has feature learning adaptability and can automatically extract features from data without the need for manual design of convolution kernels, which is in contrast to the characteristic of CNNs that rely on preset convolutional layers. ViT can process input images of various sizes without the need to crop or scale the images, while CNNs usually require fixed-size inputs, limiting their flexibility.
[0039] 2) ViT effectively captures the global information of images through the self-attention mechanism, solving the problem of insufficient global information capture that may exist in CNNs.
[0040] 3) ViT allows the understanding of the key regions of interest of the model by visualizing the attention weights, improving the interpretability of the model, while CNNs are usually insufficient in this regard.
[0041] Step S5: Process the fused features of geometry and text, and the fused features of vision and position according to the Transformer decoder to obtain the output of the Transformer decoder.
[0042] Specifically, please refer to Figure 2 , which shows the Transformer model diagram provided by the present invention. On the left side of the model architecture is the encoder module, and on the right side is the decoder module, which together constitute the core structure of Transformer. The Transformer architecture has become an advanced model for processing sequential data in deep learning, especially outstanding in solving the problem of long-range dependencies in sequences. This model is widely used in the field of natural language processing, including but not limited to machine translation and text summarization tasks. Different from the recurrent neural network that depends on the sequential processing of sequential data, Transformer adopts a modular encoder and decoder structure.
[0043] The encoder part in the Transformer core structure diagram consists of multiple encoder blocks, which are responsible for converting the input sequential data into a set of rich feature representations. Correspondingly, the decoder part is responsible for reorganizing these feature representations and generating the output sequence. Taking machine translation as an example, the encoder converts the text in the source language into feature tokens, and the decoder generates the text in the target language based on these feature tokens. The performance of Transformer is significantly affected by the model scale and the quality of the training data.
[0044] Furthermore, please refer to Figure 6 , which shows the multi-modal feature interaction Transformer diagram provided by the present invention. The cross-modal KanMiDiffusion backbone adopts the time-varying transformer decoder in MiDiffusion. The multi-head attention in the Transformer Block has better capture of boundary constraints than the conditional one-dimensional U-Net structure. At the same time, the time step is injected through the adaptive layer normalization operator, and the conditional flow input is transmitted through the multi-head cross-attention.
[0045] Furthermore, please refer to Figure 7, representing the multi - head attention structure diagram provided by the present invention. The Transformer model is based on multi - head self - attention, which enables the model to capture long - distance relationships between tokens at different positions. Specifically, is the input sequence, and this input sequence is located in the space of, where is the length of the input sequence , and is the number of hidden dimensions. Figure 7 Includes three parallel - connected normalization processes, three parallel - connected linear transformation processes, dot - product operations and scaling, normalization exponential functions, dot - product operations, concatenation, and residual connection processes.
[0046] Each self - attention head calculates the query matrix , the key matrix , and the value matrix , obtained through the linear transformation of . The mathematical expression is: ; where represents the learnable parameter, , represents the number of hidden dimensions of a head. The output of a self - attention head is the weighted sum of value vectors. The mathematical expression is: ; Among them, represents the weighted sum of N value vectors, represents the activation function, represents matrix transpose.
[0047] For a multi - head self - attention layer with heads, the final output is obtained by connecting the outputs of each self - attention head and performing linear projection calculation. The mathematical expression is: ; where represents the output of the Transformer decoder corresponding to the input vector , is the learnable parameter, represents the feature concatenation process, represents the weighted sum of N value vectors.
[0048] Step S6: Process the output of the Transformer decoder according to the second Kolmogorov - Arnold network and the third Kolmogorov - Arnold network to generate the classification distribution of semantic labels and the Gaussian mean of geometric features.
[0049] Step S7: Generate a three - dimensional scene based on the classification distribution of semantic labels and the Gaussian mean of geometric features.
[0050] Specifically, assume that each scene is in a central world frame and consists of a two - dimensional floor plan and at most objects. Use a conditional vector to represent the floor plan. Each object in the scene has a semantic label , a centroid position , a bounding box size , a rotation angle around the vertical axis , represented as: ; where represents the planar rotation group, represents the rotation angle, ranging from 0 to 360 degrees, represents matrix transpose.
[0051] All geometric features in the scene can be mapped to . The three - dimensional indoor scene is represented as . Designate the last semantic label as an empty label to handle scenes containing less than objects and define all - zero geometric properties for an "empty" object.
[0052] To verify the feasibility of the method proposed in the present invention, the present invention conducts a simulation experiment: The present invention conducts in - depth research around the 3D - FRONT dataset. The 3D - FRONT dataset, as a highly valuable synthetic indoor scene dataset, occupies a key position in the field of 3D scene synthesis research. It elaborately constructs a vast and rich virtual world of indoor scenes, which includes up to 18,797 rooms. These rooms are not randomly pieced together but are ingeniously designed and composed of 7,302 textured 3D objects carefully arranged. Each room is like a digital reproduction of a real indoor space.
[0053] The 3D-FRONT dataset is widely used in academia and industry as an important benchmark for evaluating 3D scene synthesis methods. Especially in the layout and semantic understanding of indoor environments, its role is irreplaceable. In indoor layout evaluation, researchers can use this dataset to compare and analyze the indoor layout schemes generated by different algorithms, and judge key indicators such as the rationality of the layout and space utilization. At the semantic understanding level, by mining the semantic information of the rooms in the dataset, it can help develop more intelligent indoor scene analysis systems. Each room in the 3D-FRONT dataset is equipped with extremely detailed layout and semantic information. In terms of layout, the shape, size of the room and the positional relationship of each object in space are accurately recorded. This meticulous layout information provides accurate data support for studying indoor space planning. In terms of semantic information, the dataset has detailed object annotations. Each object is assigned a semantic label closely related to it, such as "sofa", "dining table", "wardrobe", etc. These labels clarify the category and function of the object, build a key bridge for machines to understand the semantics of indoor scenes, and also provide an indispensable basis for the research of semantic-based indoor scene analysis and synthesis algorithms.
[0054] In the present invention, the proposed cross-modal KanMiDiffusion 3D scene indoor generation algorithm uses the Python language for compilation work. The Python language, with its concise syntax, rich library resources and strong scalability, provides strong support for the efficient implementation of the algorithm. In terms of model construction, it relies on the Pytorch framework.
[0055] Regarding the operating environment of the experiment, the hardware platform uses NVIDIA GeForce RTX3090. This hardware has powerful graphics processing capabilities and parallel computing performance, which can efficiently accelerate the training and inference processes of deep learning models, greatly shortening the time required for the experiment. In the software environment, the Ubuntu18.04 system is used. Its stable and reliable characteristics and good compatibility with various open-source software provide stable operating system support for the smooth development of the entire experiment. At the same time, Python3.8 version is configured. This version achieves a good balance in terms of syntax features and library support, meeting the requirements of the development and operation of this algorithm. In addition, CUDA11.3 provides efficient parallel computing acceleration capabilities for NVIDIA GPUs, giving full play to the hardware performance advantages of RTX3090. And Pytorch1.11 ensures the compatibility and efficient cooperation of the model with CUDA and other related libraries during operation.
[0056] During the optimization process, the Adam optimizer was adopted. The Adam optimizer combines the advantages of the Adagrad and RMSProp algorithms, and can adaptively adjust the learning rate of each parameter. While ensuring the convergence speed of the model, it effectively avoids the oscillation problem during the parameter update process. In terms of specific parameter settings, the learning rate was set to 0.001. This value has been debugged through multiple experiments. While ensuring that the model can effectively learn the data features, it avoids the problem that the model does not converge due to too large a learning rate or the training process being too slow due to too small a learning rate. The number of training times for the entire model was set to 500 times. Through these 500 trainings, the model can fully learn the complex features and patterns in the input data, so as to achieve effective modeling and prediction for the indoor generation task of three-dimensional scenes.
[0057] Figure 8 Figure 1 shows a three-dimensional indoor scene synthesized using the cross-modal KanMiDiffusion algorithm, which demonstrates the reconstruction effect of the three-dimensional indoor scene layout generated using the proposed cross-modal KanMiDiffusion algorithm. The algorithm aims to achieve high-quality conversion from a given two-dimensional floor plan to a three-dimensional indoor scene, and its core lies in fusing multi-modal information to enhance the accuracy and realism of scene synthesis. From Figure 8 it can be observed that the algorithm has performed precise three-dimensional reconstruction of the layout of the indoor scene. Each object in the scene, such as sofas, chairs, tables, etc., has been reasonably arranged in space according to the layout information in the two-dimensional floor plan. In addition, the algorithm can also generate three-dimensional models with high consistency according to the semantic category and size information of the objects, which is intuitively reflected in the figure. Even in the case of multiple objects and complex layouts, the algorithm can still generate a three-dimensional scene with a clear structure and reasonable layout.
[0058] Figure 9 Figure 2 shows a three-dimensional indoor scene synthesized using the cross-modal KanMiDiffusion algorithm. This scene reflects the effectiveness and accuracy of the algorithm in dealing with indoor environments with specific functional area divisions. It can be observed that the scene consists of several key areas: a rest area with a sofa and a coffee table, a dining table area, and a closet area that may be used for display or storage. The layout of these areas and the placement of objects all follow the basic principles of indoor space design, such as functional separation, smooth movement lines, and reasonable space utilization.
[0059] Figure 10 Figure 3 shows a three-dimensional indoor scene synthesized using the cross-modal KanMiDiffusion algorithm. From Figure 10In the scene, it can be observed that the scene is mainly composed of a large black area, which represents an empty room or an open space with a specific function. A small brown square is placed in a corner of the room, which symbolizes a piece of furniture or a decoration, such as a bookshelf, a side table or other small facilities. In addition, the scene also contains some small white and gray squares, representing other objects or decorative elements in the room.
[0060] The layout of this scene is simple but functional. The central position of the large black area provides sufficient space for possible furniture placement and also offers a range for the users of the room to move freely. The brown square in the corner and the small squares around it add details to the scene, making the whole space more rich and interesting. The application of the cross-modal KanMiDiffusion algorithm in this scene demonstrates its ability in dealing with minimalist-style indoor layouts. The algorithm can accurately infer the layout of the three-dimensional space and the placement of objects based on the input two-dimensional floor plan information, and can maintain the rationality and authenticity of the scene even when the number of objects is small or the space is relatively empty.
[0061] Figure 11 Four pictures of a three-dimensional indoor scene synthesized using the cross-modal KanMiDiffusion algorithm are shown. An elaborately designed indoor space can be observed, including a rest area and a possible dining area. The rest area is equipped with a sofa and a coffee table, while the dining area has a dining table and chairs. The layout of these furniture follows the principles of indoor space design, such as the division of functional areas, spatial fluidity, and symmetry in furniture placement. In addition, the size and proportional relationships of the furniture and decorative elements in the scene are precisely controlled, which visually enhances the realism and coordination of the scene. The algorithm can also generate three-dimensional models with high consistency according to the semantic categories and size information of the furniture, which is intuitively reflected in the figure.
[0062] Figure 12Shows five pictures of a 3D indoor scene synthesized using the cross-modal KanMiDiffusion algorithm. The generated layout demonstrates a high degree of rationality and harmony in spatial organization. The algorithm precisely processes object positioning, such as the visual and functional relationships between the sofa, coffee table, and TV cabinet. It can be observed that the indoor scene includes a rest area and a dining area. The rest area is equipped with a sofa and a coffee table, while the dining area has a dining table and chairs. In addition, the scene also contains some decorative elements, such as artworks and plants on the wall, which enrich the visual effect of the space and enhance the realism of the scene. The layout of the sofa and coffee table forms a comfortable rest area visually. The coffee table is located in the center of the sofa, facilitating the placement of items, which reflects the application of ergonomics in spatial layout. The placement of the dining table and chairs takes into account the convenience of dining and the openness of the space. There is enough space around the dining table for people to move and communicate. The size and proportional relationships of the furniture and decorative elements in the scene are precisely controlled, which enhances the realism and coordination of the scene visually. The algorithm can also generate 3D models with high consistency based on the semantic categories and size information of the furniture, which is intuitively reflected in the figure.
[0063] The advantage of the cross-modal KanMiDiffusion algorithm lies in its ability to handle the information fusion problem between different modalities, thereby achieving more refined and realistic effects in 3D scene synthesis. By introducing multi-modal data, the algorithm can not only learn the spatial relationships between objects but also capture the detailed features of different materials and textures, which is crucial for enhancing the visual realism of the scene.
[0064] Embodiment 2 Please refer to Figure 13 , which shows the internal structure diagram of the 3D scene automatic generation system based on scene semantics and geometric constraints provided by the present invention, including: The DeBERTa text encoder module 100 is used to perform text encoding processing on the latent variables to obtain text encoding features; The first Kolmogorov - Arnold module 200 is used to perform encoding processing on the geometric structure of the preset indoor objects to obtain geometric encoding features; The fusion feature generation module 300 is used to add the geometric encoding features and the text encoding features to generate the fusion feature of geometry and text; The Dinov2 visual encoder module 400 is used to perform visual encoding processing on the preset indoor floor plan and perform feature splicing with the preset position encoding features to obtain the fusion feature of vision and position; The Transformer encoder module 500 is used to process the fused features of geometry and text, and the fused features of vision and position to obtain the output of the Transformer decoder; The semantic label and geometric feature generation module 600 is used to process the output of the Transformer decoder according to the second Kolmogorov - Arnold network and the third Kolmogorov - Arnold network to generate the classification distribution of semantic labels and the Gaussian mean of geometric features; The three - dimensional scene generation module 700 is used to generate a three - dimensional scene based on the classification distribution of semantic labels and the Gaussian mean of geometric features.
[0065] The beneficial effects of the three - dimensional scene automatic generation system based on scene semantics and geometric constraints provided by the present invention are as follows: First, by using the DeBERTa text encoder module 100, the Transformer is used for feature extraction, the attention mechanism is fused, and the semantic relationship between words is added, so as to better represent the dynamic word vector.
[0066] Second, by using the first Kolmogorov - Arnold module 200, a learnable activation function on the edge is introduced, which fundamentally changes the neural network structure. By using the learnable one - dimensional spline function, the linear weight matrix is eliminated.
[0067] Third, by using the Transformer encoder module 500, the model can capture long - distance relationships between tokens at different positions.
[0068] The three - dimensional scene automatic generation system based on scene semantics and geometric constraints in the embodiments of the present application can be a device, or a component, an integrated circuit, or a chip in a terminal. The device can be a mobile electronic device or a non - mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, an in - vehicle electronic device, a wearable device, an ultra - mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. The non - mobile electronic device can be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self - service machine, etc. The embodiments of the present application do not make specific limitations.
[0069] The 3D scene automatic generation system based on scene semantics and geometric constraints in the embodiments of the present application can be a device with an operating system. The operating system can be the Android operating system, can be the iOS operating system, or can also be other possible operating systems, which are not specifically limited in the embodiments of the present application.
[0070] The 3D scene automatic generation system based on scene semantics and geometric constraints provided by the embodiments of the present application can implement Figures 1 to 12 each process implemented by the 3D scene automatic generation system in the method embodiments. To avoid repetition, it will not be elaborated here.
[0071] Optionally, the embodiments of the present application further provide an electronic device, including a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements each process of the above-mentioned 3D scene automatic generation method embodiments, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0072] The embodiments of the present application further provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by the processor, it implements each process of the above-mentioned 3D scene automatic generation method embodiments, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0073] Among them, the processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs, etc.
[0074] It should be noted that in the present invention, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising that element. In addition, it should be pointed out that the scope of the methods and apparatuses in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0075] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to enable a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0076] The embodiments of the present application have been described above with reference to the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.
Claims
1. A method for automatically generating a three-dimensional scene based on scene semantics and geometric constraints, characterized in that Including: Construct a preset cross-modal KanMiDiffusion denoising network model, and perform text encoding processing on the latent variables according to the DeBERTa text encoder inside the KanMiDiffusion denoising network model to obtain text encoding features. Among them, the cross-modal KanMiDiffusion denoising network model includes: a first Kolmogorov - Arnold network, a second Kolmogorov - Arnold network, a third Kolmogorov - Arnold network, a Transformer decoder, a DeBERTa text encoder, and a Dinov2 visual encoder; According to the first Kolmogorov - Arnold network, perform encoding processing on the geometric structure of the preset indoor object to obtain geometric encoding features; Add the geometric encoding features and the text encoding features to generate a geometric and text fusion feature; After performing visual encoding processing on the preset indoor floor plan according to the Dinov2 visual encoder, perform feature splicing with the preset position encoding features to obtain a visual and position fusion feature; According to the Transformer decoder, process the geometric and text fusion feature and the visual and position fusion feature to obtain the output of the Transformer decoder; According to the second Kolmogorov - Arnold network and the third Kolmogorov - Arnold network, process the output of the Transformer decoder to generate a classification distribution of semantic labels and a Gaussian mean of geometric features; Generate a 3D scene based on the classification distribution of semantic labels and the Gaussian mean of geometric features.
2. The three-dimensional scene automatic generation method based on scene semantics and geometric constraints according to claim 1, characterized in that The Kolmogorov - Arnold network includes multiple KAN Layer layers, and the mathematical expression is: ; Among them, represents the th layer of the KAN network, represents the th layer of the KAN network, represents the first layer of the KAN network, represents the 0th layer of the KAN network, represents the input vector, represents the input vector corresponding output result; Among them, each KAN network has dimensional input and dimensional output. represents multiplication processing, which is composed of a learnable activation function and the mathematical expression is: ; The mathematical expression of the Kolmogorov - Arnold network from layer to layer is as follows: ; Among them, represents the output vector of the th layer, represents the output vector of the th layer, represents the th layer matrix of the KAN network.
3. The three-dimensional scene automatic generation method based on scene semantics and geometric constraints according to claim 1, wherein The DeBERTa text encoder includes: multiple layers of Transformer encoder layers connected in parallel, a word embedding layer, and a word vector representation layer. Each layer of the Transformer encoder layer includes multiple Transformer encoders connected in series. Among them, the word embedding layer is used to embed each position information, sentence embedding, and word embedding as input vectors, and input them into the multiple layers of Transformer encoder layers connected in parallel for encoding processing to generate word vectors corresponding to each Transformer encoder layer.
4. The three-dimensional scene automatic generation method based on scene semantics and geometric constraints according to claim 1, wherein The Dinov2 visual encoder is based on ViT, divides the image into multiple small blocks to obtain an image block sequence to capture long-range dependencies in the image. The Dinov2 visual encoder includes multiple layers of Transformer modules to generate high-level feature representations rich in semantics.
5. The three-dimensional scene automatic generation method based on scene semantics and geometric constraints according to claim 1, wherein The Transformer decoder includes a plurality of Transformer blocks. Each Transformer block is sequentially connected to a first adaptive normalization layer, a multi-head self-attention layer, a second adaptive normalization layer, a multi-head cross-attention layer, and a feed-forward neural network layer. The output of the Transformer decoder is obtained by concatenating the outputs of each self-attention head in the multi-head self-attention layer and performing a linear projection calculation. The mathematical expression of the output of the Transformer decoder is: ; Among them, represents the input vector corresponding to the output of the Transformer decoder, are learnable parameters, represents the feature concatenation process, represents the weighted sum of N value vectors.
6. A three-dimensional scene automatic generation system based on scene semantics and geometric constraints, characterized in that, Including: A DeBERTa text encoder module for performing text encoding processing on latent variables to obtain text encoding features; A first Kolmogorov - Arnold module for encoding the geometric structure of a preset indoor object to obtain geometric encoding features; A fusion feature generation module for adding the geometric encoding features and the text encoding features to generate a geometric and text fusion feature; A Dinov2 visual encoder module for performing visual encoding processing on a preset indoor floor plan and then concatenating features with preset position encoding features to obtain a visual and position fusion feature; A Transformer encoder module for processing the geometric and text fusion feature and the visual and position fusion feature to obtain the output of the Transformer decoder; A semantic label and geometric feature generation module for processing the output of the Transformer decoder according to a second Kolmogorov - Arnold network and a third Kolmogorov - Arnold network to generate a classification distribution of semantic labels and a Gaussian mean of geometric features; A three-dimensional scene generation module for generating a three-dimensional scene based on the classification distribution of semantic labels and the Gaussian mean of geometric features.
7. An electronic device, characterized in that, Including a processor, a memory, and a program or instructions stored on the memory and executable on the processor. When the program or instructions are executed by the processor, the steps of the three-dimensional scene automatic generation method based on scene semantics and geometric constraints as described in any one of claims 1 - 5 are implemented.
8. A readable storage medium, characterized in that, A program or instructions are stored on the readable storage medium. When the program or instructions are executed by the processor, the steps of the three-dimensional scene automatic generation method based on scene semantics and geometric constraints as described in any one of claims 1 - 5 are implemented.
Citation Information
Patent Citations
Mobile robot positioning method and system based on three-dimensional point cloud and vision fusion
CN111429574A
Mouse characteristic behavior analysis method based on scene geometric constraint and deep learning
CN114724057A
Traffic scene generation type image description method
CN117173450A
Layout-controllable three-dimensional scene characterization and generation method based on large language model
CN117409140A
Scene text recognition method and device based on diffusion model and readable medium
CN117911997A
Cited By
Mixed diffusion model-based workshop equipment layout generation method and equipment, and storage medium
CN120509104A