Desktop layer object generation method based on conditional diffusion model
Through the desktop object generation method (DESK-GEN) based on the conditional diffusion model, desktop style and geometric features are extracted to generate object layouts that conform to functional and physical rules, solving the problem of insufficient desktop object placement in existing technologies and improving the realism of the virtual environment.
Patent Information
- Application Number
- CN202511164098.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-08-20
AI Technical Summary
Existing technologies ignore the placement of objects on desktop furniture in indoor three-dimensional scene synthesis datasets, resulting in insufficient realism in the virtual environment. There is a relative lack of research on the application of generative artificial intelligence technology in the field of desktop object generation.
A desktop object generation method (DESK-GEN) based on the conditional diffusion model is adopted. By extracting desktop style features and geometric features, combined with a U-shaped network, an object configuration that meets functional requirements and has a reasonable layout is generated. The semantic and geometric condition vectors are extracted using preprocessed datasets and pre-trained models, and the denoising network is driven to predict object categories and pose parameters.
The spatial compliance and detail expression of object layout are significantly improved, while ensuring the uniformity of desktop scene style. The generated object layout conforms to physical rules and style requirements.
Smart Images

Figure CN120655856A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to three-dimensional reconstruction of computer vision, and in particular to a method for generating object placement of desktop furniture. Background Art
[0002] The desktop object generation task is a subtask of indoor 3D scene layout generation and is of great significance in the field of computer vision. It involves not only the position, orientation, and size of desktop objects, but also their categories and relative relationships. This task is particularly critical in application scenarios such as game development, virtual reality, and augmented reality, enhancing the realism of virtual environments. However, most synthetic indoor 3D scene datasets ignore the placement of objects on desktop furniture.
[0003] In recent years, generative artificial intelligence technology has flourished, and a variety of representative generation paradigms have emerged, including autoregressive-based sequence generation methods, adversarial generation methods based on generative adversarial networks (GANs), and probabilistic generation methods based on diffusion models. These methods have demonstrated excellent performance in tasks such as image synthesis and scene generation. However, application research in the field of desktop object generation is relatively scarce. Summary of the Invention
[0004] In response to the above problems, the present invention proposes a desktop layer object generation method (DESK-GEN) based on a conditional diffusion model. The goal of this algorithm is to generate a desktop layer object configuration that meets functional requirements and has a reasonable layout by inputting two-dimensional image data and three-dimensional point cloud data containing desktop layer furniture. Specifically, the DESK-GEN method proposed in the present invention first extracts the material, shape and style from the two-dimensional image data of desktop layer furniture through desktop style features to generate a semantic condition vector; at the same time, the desktop geometric feature extraction module is used to process the point cloud to construct a geometric condition vector containing spatial distribution and size features. In the diffusion model generation stage, the semantic condition vector and the geometric condition vector of spatial distribution and size features are jointly encoded through the desktop layer object generation module, driving the denoising network to predict the category distribution and posture parameters of the object, and finally outputting a set of objects that meet functional consistency and physical rationality. While ensuring the uniformity of the desktop scene style, this method significantly improves the spatial compliance and detail expression of the object layout.
[0005] Several key parts of the present invention are described as follows:
[0006] 1. Adjust the training and validation datasets to obtain desktop scene data, including images from different viewpoints and furniture point cloud data, and preprocess them to achieve a unified standard. The datasets used in this method are desktop scene data from three types of environments: residential, office, and commercial. Dimensions include, but are not limited to: dense point cloud data of different furniture scenes, high-definition images from different viewpoints, object categories, 3D bounding box annotations, and multi-level semantic systems such as material type, functional attributes, and style labels. Different datasets will have different dimensions, and they need to be unified to this consistent standard before training and inference.
[0007] 2. Desktop Style Extraction: Utilizing images from different perspectives and their corresponding semantic labels in the pre-processed training dataset, a pre-trained model is used to extract desktop style semantic features, guiding the uniformity of object category generation.
[0008] 3. Desktop Geometric Style Extraction: Utilizing the pre-processed furniture point cloud data from the training dataset, we extract geometric features and constraint features from the desktop furniture point cloud data. Specifically, we extract the surface point cloud Xsurface and the boundary point cloud Xcontour to control the placement of generated objects.
[0009] 4. Desktop object generation: Integrate desktop style semantic features, surface point cloud Xsurface and boundary point cloud Xcontour, use U-net to build a diffusion model, and train a model to generate the desktop layer.
[0010] 5. Predictive reasoning for desktop generation: Using the diffusion model trained in the previous step, we select different scenarios, such as desks, coffee tables, and nightstands, and infer and generate desktop layouts for corresponding scenarios, achieving desktop object layout generation that conforms to physical rules and style requirements.
[0011] Beneficial effects and innovations of the present invention:
[0012] 1. Preprocessing methods for training and validation data, unifying standards and classifications, and providing feasible and efficient methods for subsequent processing of different scenarios.
[0013] 2. A comprehensive and effective desktop graphics generation method, including desktop style extraction based on a pre-trained model, desktop geometric feature extraction using an autoencoder, and finally, denoising layout generation using a U-Net. This method accurately generates corresponding object arrangements based on different desktop scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 It is the overall flow chart of the method of the present invention;
[0015] Figure 2Generate a denoising diffusion network structure U-net for desktop layer objects;
[0016] Figure 3 This is a schematic diagram of the desktop style feature extraction module structure;
[0017] Figure 4 This is the desktop point cloud data processing process;
[0018] Figure 5 This is an example of desktop point cloud processing results;
[0019] Figure 6 Schematic diagram of the results generated for the present invention. DETAILED DESCRIPTION
[0020] like Figure 1 As shown, a method for generating desktop layer objects based on a conditional diffusion model specifically includes the following steps:
[0021] Step 1: Preprocessing of training and validation datasets:
[0022] The specific preprocessing methods are: unifying the classification standards of data sets, screening data, and specifying the classification range and coordinate units.
[0023] Normalize the geometric parameters of the point cloud. To address the discreteness of the model size, scale normalization is performed based on the diagonal length of the object's 3D bounding box. The model vertex coordinates are scaled to the unit bounding box space to eliminate size deviations caused by modeling scale differences.
[0024] Orientation unification: Define the main orientation axis according to the object's function (such as the positive Z axis of the display screen), and align the model's local coordinate system with the global coordinate system through rotation matrix transformation to ensure that the object's pose parameters during the generation process are consistent with the real-world semantics.
[0025] Data augmentation: geometric perturbations are applied to the desktop point cloud, including random rotation transformation (θ∈[−15°, 15°]), translation perturbation (Δx, Δy∈[−0.2m, 0.2m]), and scale scaling (s ∈ [0.9, 1.1]) to improve the model's robustness to spatial transformations.
[0026] Step 2: Desktop style feature extraction:
[0027] This step uses a pre-trained model (including but not limited to BLIP-2) as the base model to extract the style category information of the desktop and convert it into a conditional vector that can be used in the diffusion model. The style of desktop furniture can be defined as features related to the application scenario or functional category, such as coffee table, desk, etc. These categories not only reflect the diversity of desktop scenes, but also provide style classification and semantic guidance for the generation process. Different categories of desktop furniture have significant differences in materials, shapes, placement habits, etc. Therefore, in the desktop object generation task, accurately extracting and encoding these category information is crucial to ensure style consistency.
[0028] like Figure 3 As shown, the desktop style semantic feature extraction function consists of an image encoder, a vision-language translator (Q-Former), and a large language model (LLM) decoder. This function extracts rich desktop style information through vision-language modeling. First, a pre-trained image encoder extracts low-level visual features from a 2D image of the desktop furniture. This encoder effectively captures the object's shape, texture, material, and scene layout. These visual features are then fed into the Q-Former for cross-modal processing. The Q-Former, employing a Transformer architecture, learns high-level visual-semantic alignments through a self-attention mechanism, enabling different desktop styles to be mapped into a unified embedding space. In this process, the Q-Former generates a set of learnable embedding sequences that not only capture desktop categories (such as desk, coffee table, kitchen table, etc.), but also implicitly encode environmental features such as the layout of items on the desktop, color style, and functional attributes.
[0029] The system achieves cross-modal generation through a cascaded architecture of "visual encoding → language decoding", specifically as follows: Q-Former first compresses pixel-level image information into a compact visual semantic representation; these visual representations serve as conditional inputs to the LLM; and the LLM generates the final style description text (T) based on the visual semantic context.
[0030] This feature can be used to control the subsequent diffusion model, enabling it to be personalized according to the desktop style.
[0031] In addition, the learnable embedding mechanism of this feature enables the model to adapt to different desktop environments without manually defining desktop categories, which improves the flexibility of style expression.
[0032] Step 3: Desktop geometric style extraction:
[0033] Input desktop scene point cloud data X ∈ R N×3(Where N is the number of point clouds, and the three-dimensional coordinates include x, y, and z-axis spatial information) The semantic labels provided by the desktop scene dataset are used to filter instance objects belonging to objects on the desktop, and the bounding box parameter range is calculated based on the object instance name index.
[0034] In order to extract the desktop layer from the point cloud data, it is necessary to first determine the height range of the desktop furniture. The specific steps are as follows:
[0035] 1. Calculate the desktop height baseline: First, obtain all coordinate values of the desktop point cloud in the vertical direction (Z axis), denoted as set Z. Take the minimum value in Z as the reference baseline, and add a preset thickness compensation value δ (δ ranges from 0.05 to 0.1 meters) to obtain the final support surface baseline height hbase.
[0036] 2. Surface point cloud extraction: Based on the base height hbase and δ, the point cloud data with a height within the interval [hbase − δ, hbase] is filtered out. This part of the point cloud is the surface point cloud Xsurface of the desktop.
[0037] 3. Boundary point cloud extraction: To further extract the boundary contour of the desktop, the following method is used:
[0038] Horizontal plane projection: Project the surface point cloud Xsurface onto the horizontal plane (ignoring the Z coordinate) to obtain a two-dimensional point set.
[0039] Convex hull calculation: Perform convex hull calculation on the projected two-dimensional point set to find the smallest convex polygon that can contain all points. Its vertex set is recorded as V.
[0040] 3D boundary reconstruction: Remap the 2D coordinates of the convex hull vertex V back to 3D space, that is, set the Z coordinate of each vertex vj to the base height hbase, and thus obtain the 3D boundary point cloud Xcontour.
[0041] Through the above steps, the surface point cloud Xsurface and boundary point cloud Xcontour of the desktop can be separated from the original point cloud data.
[0042] In the data normalization stage, affine transformation is used to eliminate scale differences: the extracted surface point cloud (Xsurface) and boundary point cloud (Xcontour) are centered respectively X = X − µ, where µ is the coordinate of the point cloud centroid. At the same time, the centered point cloud is scaled to standardize its distribution range:
[0043] Calculate the maximum Euclidean distance of all point cloud coordinates (i.e. the distance of the point farthest from the origin).
[0044] Divide each point cloud by this maximum distance to ensure that all points fall within the unit sphere (i.e., their distance from the origin does not exceed 1).
[0045] This processing makes desktop point clouds of different sizes comparable in feature space, laying the foundation for feature decoupling learning of the subsequent VAE encoder, and finally outputting the corresponding standardized point cloud.
[0046] The normalized point cloud is then fed into the feature encoding module. The encoder network consists of multiple layers of graph convolution (GraphConv) modules. The decoder of this point cloud reconstruction model uses an inverse graph convolution layer to reconstruct the point cloud coordinates X (in N×3 dimensional real space).
[0047] The reconstruction loss function uses the chamfer distance, and its calculation process includes a point-by-point distance operation between the original point cloud X and the estimated value X. The specific implementation is: first calculate the sum of the squares of the shortest distances from each point in the two sets of point clouds to the other set of point clouds, and then average the two sets of results and add them together.
[0048] To maintain the regularity of the latent space, the model introduces a KL divergence loss term. This is calculated by summing all latent dimensions: the calculation of each dimension includes the logarithmic variance, the mean square, and the variance adjustment term, and the final result is half.
[0049] The total loss function is the weighted sum of the chamfer distance loss and the KL divergence loss, where the loss weight coefficient β adopts an annealing strategy - it is set to 0.01 at the beginning of training and gradually increases linearly to 1.0 during training to balance the reconstruction accuracy and the strength of distribution regularization.
[0050] The network finally outputs two features: surface point cloud features and boundary point cloud features. These features will serve as conditional inputs for the subsequent diffusion model denoising process. The final generated effect can be referred to Figure 5 The results are shown in the following figure. The first figure is a schematic diagram of the original state of the dataset, the second figure is a schematic diagram of the desktop layer area, and the third figure is the desktop point cloud extraction result. The upper part is the surface point cloud style, and the lower part is the boundary point cloud style.
[0051] Step 4: Desktop object generation
[0052] This module generates desktop object layouts that conform to both physical rules and style requirements by integrating desktop style features, desktop point cloud geometry, and boundary contour constraints. The following sections will analyze each of these conditional features and elaborate on the module's conditional fusion and generation control during the denoising process.
[0053] 4.1 Conditional Feature Analysis
[0054] This algorithm guides the generation process through a multi-conditional control mechanism, and the three types of conditional features work together to ensure that the generated layout meets geometric constraints and semantic adaptability.
[0055] Desktop style semantic features: A pre-trained model performs semantic parsing on the input 2D image, generating a 256-dimensional feature vector representing the functional attributes of the desktop furniture through vision-language alignment (e.g., "desk" corresponds to a preference for stationery objects, and "dining table" is associated with the distribution of tableware objects). This feature vector is mapped into a conditional embedding vector via a fully connected network. The channel-wise attention mechanism interacts with the diffusion network features to guide the generation of functional consistency between object categories and desktop types.
[0056] Desktop point cloud geometry features: Figure 4 In [1], the 32-dimensional latent vector extracted by the VAE encoder from the surface point cloud encodes the geometric properties of the desktop support surface. These geometric features serve as control conditions for the diffusion model, ensuring that the generated object layout conforms to spatial constraints. Furthermore, these features provide a global representation of the desktop surface shape, enabling the generated model to better match the desktop's form during layout, improving rationality and physical consistency.
[0057] Boundary Contour Constraints: A 32-dimensional feature vector generated by VAE encoding the boundary point cloud represents the three-dimensional spatial constraints of the desktop boundary. This helps objects on the desktop layer perceive the desktop's boundaries, allowing them to be placed appropriately within the valid area, preventing them from appearing outside the desktop or overlapping the edge. This condition effectively improves the placement accuracy of desktop objects, making the generated results more consistent with the desktop shape and avoiding unreasonable layouts caused by missing boundary information.
[0058] Three types of conditions form a coordinated control system of "semantics, geometry, and boundaries": style semantic features control the functional rationality of object categories, point cloud geometric features optimize the contact relationship between objects and support surfaces, and boundary contour features impose hard spatial constraints. This multi-conditional feature fusion mechanism ensures that the generation process adheres to functional logic, physical rules, and spatial boundaries simultaneously, providing the core driving condition for achieving high-fidelity generation.
[0059] 4.2 Conditional Feature Fusion and Generative Control
[0060] This module addresses the issues of style consistency and layout rationality of objects in complex desktop scenes, and proposes a generation control mechanism for hierarchical conditional injection. Through multi-conditional feature fusion and attention guidance strategies, it achieves controllable optimization of the diffusion model generation process. The specific process is as follows:
[0061] First, the three types of conditional features need to be preprocessed. The 256-dimensional style semantic features extracted by the pre-trained model BLIP-2 (characterizing the desktop functional type), the 32-dimensional geometric features encoded by the surface point cloud VAE (describing the geometric shape of the support surface), and the 32-dimensional boundary features encoded by the boundary point cloud VAE (defining the spatial range of the effective area) are channel-concatenated to form a 320-dimensional joint conditional vector. This vector is compressed to 128 dimensions through a two-layer fully connected network to obtain the conditional code C ∈ R 128 .
[0062] Later in Figure 2 In the architecture design of the diffusion network U-Net shown in Fig. 1, a cross-attention layer is embedded after each residual block to establish a dynamic interaction mechanism between the noise latent code and the conditional code. Specifically, the query vector Q ∈ R H×W×d The spatial feature map derived from the current noise latent code (where H, W are the feature map sizes and d = 128 is the feature dimension), the key-value pair K, V ∈ R 128×d Then the conditional code C is learned by the linear projection matrix Wk, Wv∈R 128×128 generate.
[0063] The calculation process of attention weights follows the standard scaled dot product attention formula:
[0064]
[0065] The temperature coefficient (Here d = 128, the actual calculated value is 11.31) is used to stabilize gradient propagation and prevent the Softmax function from entering the saturation region.
[0066] Figure 6 This is a schematic diagram of the results generated by the present invention. It can be seen that this method allows the generation process to maintain the creativity of the diffusion model while strictly following the physical space constraints and functional logic, and ultimately outputs a high-quality layout that meets expectations.
Claims
1. A method for generating desktop objects based on a conditional diffusion model, characterized in that: The following processes are included: Acquire desktop scene data, including images from different perspectives and furniture point cloud data, and perform preprocessing; Using images from different perspectives and corresponding semantic labels in the pre-processed training dataset, the pre-trained model is used to extract desktop style semantic features. Using the furniture point cloud data in the preprocessed training dataset, the geometric features and constraint features are extracted from the desktop furniture point cloud data to obtain the surface point cloud and boundary point cloud. A diffusion model is constructed using a U-shaped network, which integrates desktop style semantic features, surface point cloud, and boundary point cloud to obtain a model for generating the desktop layer. Using the diffusion model trained in the previous step, we select different scenes, infer and generate the desktop layout in the corresponding scenes, and realize the generation of desktop objects.
2. The method for generating desktop objects based on the conditional diffusion model according to claim 1, characterized in that: The specific implementation process of extracting desktop style semantic features is as follows: The pre-trained model is used as the base model to extract the style category information of the desktop and convert it into a conditional vector that can be used in the diffusion model. The style of desktop furniture is defined as features related to the application scenario or functional category. The desktop style semantic feature extraction function consists of an image encoder, a visual-language translation Q-Former, and a large language model LLM decoder, which extracts desktop style information through visual-language modeling.
3. The method for generating desktop objects based on the conditional diffusion model according to claim 2, characterized in that: The desktop style semantic feature extraction function is specifically implemented as follows: First, the input 2D image of the tabletop furniture is passed through a pre-trained image encoder to extract the underlying visual features. The encoder can capture the object's shape, texture, material, and scene layout information. The visual features are then fed into a Q-Former for cross-modal processing. Using a Transformer architecture, Q-Former uses a self-attention mechanism to learn visual-semantic alignment, enabling different desktop styles to be mapped into a unified embedding space. In this process, Q-Former generates a set of learnable embedding sequences that not only capture the desktop categories but also implicitly encode environmental features. The embedding sequence is input into the LLM as a condition, and the LLM generates the final style description text, namely the desktop style semantic features, based on the visual semantic context.
4. The method for generating desktop objects based on the conditional diffusion model according to claim 3, characterized in that: The specific implementation process of extracting geometric features and constraint features is as follows: Get all the vertical coordinates of the desktop point cloud, recorded as set Z; take the minimum value in Z as the reference base, and add a preset thickness compensation value δ to obtain the final support surface base height hbase; According to the base height hbase and δ, the point cloud data with a height within the interval [hbase−δ, hbase] is filtered out. This part of the point cloud is the surface point cloud Xsurface of the desktop; Project the surface point cloud Xsurface onto the horizontal plane to obtain a two-dimensional point set. Perform convex hull calculation on the projected two-dimensional point set to find the minimum convex polygon that can contain all points. Its vertex set is recorded as V. Remap the two-dimensional coordinates of the convex hull vertex V back to three-dimensional space, that is, set the Z coordinate of each vertex vj to the base height hbase, and obtain the three-dimensional boundary point cloud Xcontour.
5. The method for generating desktop objects based on the conditional diffusion model according to claim 4, characterized in that: The process of extracting geometric features and constraint features also includes: In the data normalization stage, affine transformation is used to eliminate scale differences: the extracted surface point cloud Xsurface and boundary point cloud Xcontour are centered separately, and the centered point clouds are scaled to standardize their distribution range: the maximum Euclidean distance of all point cloud coordinates is calculated; each point cloud coordinate is divided by the maximum Euclidean distance, and the corresponding standardized point cloud is finally output; The normalized point cloud is then input into the feature encoding module; the encoder network consists of a multi-layer graph convolution GraphConv module; the decoder uses the inverse graph convolution layer to reconstruct the point cloud coordinates to obtain the estimated value X; The reconstruction loss function uses the chamfer distance loss, which is calculated by performing a point-by-point distance operation on the original point cloud X and the estimated value X. The specific implementation is as follows: first, the sum of the squares of the shortest distances from each point in the two point clouds to the other point cloud is calculated, and then the two sets of results are averaged and added together; The KL divergence loss term is introduced, which is calculated by summing all potential dimensions: the calculation of each dimension includes the logarithmic variance, mean square and variance adjustment term, and the final result is half; The total loss function is the weighted sum of the chamfer distance loss and the KL divergence loss, where the loss weight coefficient adopts an annealing strategy to balance the reconstruction accuracy and the strength of distribution regularization.
6. The method for generating desktop objects based on the conditional diffusion model according to claim 5, characterized in that: The specific implementation process of the model for generating the desktop layer is as follows: The generation process is guided by a multi-condition control mechanism, and the three types of conditional features work together to ensure that the generated layout meets geometric constraints and semantic adaptability. A generation control mechanism of hierarchical conditional injection is proposed, which realizes the controllability optimization of the diffusion model generation process through the fusion of three types of conditional features and attention guidance strategy.
7. The method for generating desktop objects based on the conditional diffusion model according to claim 6, characterized in that: The three types of conditional features are: desktop style semantic features, desktop point cloud geometric features and boundary contour constraint features.
8. The method for generating desktop objects based on the conditional diffusion model according to claim 7, characterized in that: The generation control mechanism of hierarchical conditional injection achieves controllable optimization of the diffusion model generation process through multi-conditional feature fusion and attention guidance strategy. The specific process is as follows: First, the three types of conditional features are preprocessed. The style semantic features extracted by the pre-trained model, the geometric features encoded by the surface point cloud VAE, and the boundary features encoded by the boundary point cloud VAE are channel-concatenated to form a joint conditional vector. This vector is compressed through a fully connected network to obtain the conditional code C. Subsequently, in the architecture design of the diffusion network U-network, a cross-attention layer is embedded after each residual block to establish a dynamic interaction mechanism between the noise latent code and the conditional code. Specifically, the query vector Q is derived from the spatial feature map of the current noise latent code, and the key-value pairs K, V are generated by the conditional code C through a learnable linear projection matrix. The calculation process of the attention weight follows the scaled dot product attention formula.
Citation Information
Patent Citations
Method for automatically extracting elements of high-precision map
CN115588178A
Indoor point cloud scene semantic segmentation method based on patch context features
CN115620287A
Three-dimensional object generation method based on diffusion model and semantic guidance
CN116721200A
Scene-level synthetic point cloud enhanced semantic segmentation method and system based on diffusion model
CN119992082A
Layout-controllable image personalized generation method based on diffusion model
CN120014117A