Boundary representation model generation method and system based on structured hidden space, and medium

By encoding and decoding the geometric and topological information of the B-rep model in the structured hidden space, a topologically consistent boundary representation model is generated, which solves the problems of model incoherence and missing topological information in the prior art, and improves the generation efficiency and effectiveness.

CN120493324AActive Publication Date: 2025-08-15SHENZHEN UNIV

Patent Information

Application Number
CN202510363432.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-08-15
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

Existing B-rep model generation methods fail to learn geometric features and topological relationships in a unified representation space, resulting in incoherence of the generated models, missing topological information, or requiring additional post-processing steps to fix the error.

Method used

The boundary representation model generation method of structured hidden space is adopted. By encoding continuous geometric information and discrete topological information of different primitive types, a unified hidden vector is formed, and the hidden vector is decoded to restore the relationship between the primitive type and its geometric information and topological information, and a boundary representation model is generated based on the hidden vector.

Benefits of technology

Geometric features and topological information are encoded simultaneously in a unified structured hidden space to ensure the topological consistency of the generated boundary representation model, improve generation efficiency, reduce training complexity, and improve effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120493324A_ABST
    Figure CN120493324A_ABST
Patent Text Reader

Abstract

The invention discloses a structured hidden space-based boundary representation model generation method and system and a medium, and the method comprises the steps: coding continuous geometric information from different primitive types and corresponding discrete topological information, and forming a unified hidden vector with expressive power, the primitive types including curved surfaces, curves and points; decoding the implicit vectors in the structured implicit space to recover all primitive types and the relationship between the geometric information and the topological information of the primitive types; and training a hidden space diffusion model based on the hidden vector, and generating a boundary representation model based on the hidden space diffusion model according to different input conditions. According to the method, geometric features and topological information can be coded at the same time in the unified structured hidden space, the topological consistency of the generated boundary representation model is ensured, the effectiveness is improved, the efficiency of generating the boundary representation model is effectively improved, and the training complexity is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer graphics processing, and in particular to a method, system and medium for generating a boundary representation model based on a structured latent space. Background Art

[0002] B-rep (Boundary Representation) models are the most basic shape representation format for CAD models and are widely used in design, engineering, robotics, e-commerce, and other fields. However, because B-rep models contain both continuous geometric parameters (such as surfaces and curves) and discrete topological relationships, deep learning-based B-rep model generation faces many challenges.

[0003] However, the main drawback of existing B-rep model generation methods is that they fail to simultaneously learn the geometric features and topological relationships of the B-rep model in a unified representation space, resulting in incoherent generated models, missing topological information, or requiring additional post-processing steps to fix errors.

[0004] Therefore, the prior art still has defects. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to address the above-mentioned defects of the prior art and provide a method, system and medium for generating a boundary representation model based on a structured latent space. The technical solution adopted by the present invention is as follows:

[0006] In a first aspect, the present invention provides a method for generating a boundary representation model based on a structured latent space, wherein the method comprises:

[0007] Encodes continuous geometric information from different primitive types, including surfaces, curves, and points, and the corresponding discrete topological information into a unified and expressive latent vector.

[0008] Decode the latent vectors in the structured latent space to recover the relationships between all primitive types and their geometric and topological information;

[0009] A latent space diffusion model is trained based on the latent vector, and a boundary representation model is generated based on the latent space diffusion model according to different input conditions.

[0010] In one implementation, encoding continuous geometric information from different primitive types and corresponding discrete topological information to form a unified and expressive latent vector includes:

[0011] Based on a preset variational autoencoder, the geometric information of different primitive types and the corresponding discrete topological information are encoded into the surface latent space to obtain the latent vector.

[0012] In one implementation, encoding geometric information of different primitive types and corresponding discrete topological information into a surface latent space based on a preset variational autoencoder to obtain the latent vector includes:

[0013] Determine several sampling points of the surface and several curves;

[0014] Train a graph neural network to propagate curve features based on topological connectivity to surface primitives;

[0015] Use a series of self-attention layers to aggregate and exchange features in the latent space of each surface;

[0016] A multi-layer perceptron is used to map the topologically aware surface latent space into the mean and variance of the Gaussian latent space, and the latent vector is sampled from the predicted mean and variance.

[0017] In one implementation, decoding the latent vectors in the structured latent space to recover the relationships between all primitive types and their geometric and topological information includes:

[0018] Based on the preset neural intersection module, it recovers the relationship between the geometric information and topological information of curves and points based on the surface latent vector, and predicts whether two surfaces intersect;

[0019] The sampled latent vectors and intersection curve features are determined, and a decoder network consisting of CNN and upsampling layers is used to reconstruct surface primitives and curve primitives.

[0020] In one implementation, the method of restoring the relationship between the geometric information and topological information of the curve and point based on the surface latent vector based on the preset neural intersection module and predicting whether two surfaces intersect includes:

[0021] Use a series of self-attention layers to perform deep feature exchange on different surface latent vectors sampled from the structured latent space;

[0022] A series of cross-attention layers are used to exchange the features of two surfaces, where the latent vector of the first surface is used as the query vector and the latent vector of the second surface is used as the key vector and value vector;

[0023] A multi-layer perceptron layer is used to map the swapped features to the corresponding curve features, and a binary classifier is trained to determine whether the two surfaces intersect.

[0024] In one implementation, training a latent space diffusion model based on latent vectors, and generating a boundary representation model based on the latent space diffusion model according to different input conditions, includes:

[0025] Pad the latent vector to a fixed length by randomly repeating it until a predefined maximum number of surface primitives is reached;

[0026] Using the conditional vector as a constraint, the random noise is denoised and mapped into a target latent vector, wherein the conditional vector is a feature vector generated according to the input condition;

[0027] The latent space diffusion model is trained using a denoising diffusion probability model linear scheduler, and the loss function is set to the L2 loss between the target latent vector and the true value latent vector.

[0028] In one implementation, the condition vector is used as a constraint, including:

[0029] If there is no input condition, the condition vector is set to zero;

[0030] If the input condition is single view, multi-view, or sketch, for each input image, a 1024-dimensional feature vector is extracted using the pre-trained DINOv2 model. When there are multiple input images, the view pose information is embedded by adding position encoding, and the feature vectors of all images are mean-fused. The MLP layer is used to map the feature vectors to a 256-dimensional conditional vector.

[0031] If the input condition is a sparse point cloud or a dense point cloud, a 1024-dimensional feature vector is extracted from the input point cloud data using the PointNet network, and then mapped to a 256-dimensional conditional vector through the MLP layer.

[0032] In a second aspect, an embodiment of the present invention further provides a system for generating a boundary representation model based on a structured latent space, wherein the system is used to implement the steps of the method for generating a boundary representation model based on a structured latent space in any one of the above-mentioned solutions, and the system includes:

[0033] A latent vector encoding module, which encodes continuous geometric information from different primitive types, including surfaces, curves, and points, and the corresponding discrete topological information, to form a unified and expressive latent vector;

[0034] The latent vector decoding module is used to decode the latent vectors in the structured latent space to recover the relationship between all primitive types and their geometric and topological information;

[0035] The boundary representation model generation module is used to train a latent space diffusion model based on latent vectors, and generate a boundary representation model based on the latent space diffusion model according to different input conditions.

[0036] In a third aspect, an embodiment of the present invention further provides a terminal, wherein the terminal includes a memory, a processor, and a boundary representation model generation program based on structured latent space stored in the memory and runnable on the processor. When the processor executes the boundary representation model generation program based on structured latent space, the steps of the boundary representation model generation method based on structured latent space of any one of the above-mentioned schemes are implemented.

[0037] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein a boundary representation model generation program based on a structured latent space is stored on the computer-readable storage medium. When the boundary representation model generation program based on a structured latent space is executed by a processor, the steps of the boundary representation model generation method based on a structured latent space described in any one of the above-mentioned schemes are implemented.

[0038] Beneficial effects: Compared with the prior art, the present invention provides a method for generating a boundary representation model based on a structured latent space. The present invention first encodes continuous geometric information from different primitive types and corresponding discrete topological information to form a unified and expressive latent vector, wherein the primitive types include surfaces, curves, and points. Then, the latent vectors in the structured latent space are decoded to restore the relationship between all primitive types and their geometric information and topological information. Next, a latent space diffusion model is trained based on the latent vectors, and a boundary representation model is generated based on the latent space diffusion model according to different input conditions. The present invention can simultaneously encode geometric features and topological information in a unified structured latent space, ensure the topological consistency of the generated boundary representation model, improve effectiveness, effectively improve the efficiency of boundary representation model generation, and reduce training complexity. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 The present invention provides a flowchart of a preferred embodiment of a method for generating a boundary representation model based on a structured latent space.

[0040] Figure 2 A schematic diagram of the encoding process of a variational autoencoder in a boundary representation model generation method based on a structured latent space provided in an embodiment of the present invention.

[0041] Figure 3 A schematic diagram of a method for generating a boundary representation model based on a structured latent space provided in an embodiment of the present invention, which uses a neural intersection module to recover the geometric and topological relationship between curves and points based on surface latent vectors.

[0042] Figure 4 A schematic diagram of the architecture of a device for generating a boundary representation model based on a structured latent space provided by an embodiment of the present invention.

[0043] Figure 5This is a functional block diagram of a terminal provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0044] In order to make the purpose, technical solution and effect of the present invention clearer and more specific, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0045] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents, operations, or steps, nor must they be executed in the order described. For example, some operations or steps may be decomposed, combined, or partially merged, so the actual execution order may vary depending on the actual situation.

[0046] It should be understood that the terms used in this specification are only for the purpose of describing particular embodiments and are not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0047] It should be understood that, to facilitate a clear description of the technical solutions of the embodiments of the present invention, in the embodiments of the present invention, terms such as "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. For example, the first control information and the second control information are merely used to distinguish different control information and do not limit their order.

[0048] Those skilled in the art can understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit them to be different.

[0049] It should be further understood that the term "and / or" used in the present description and the appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0050] The existing B-rep model generation methods mainly include the following methods:

[0051] 1. Step-by-step training methods: For example, BRepGen (a diffusion-based generation method) and SolidGen (a Transformer and pointer network-based generation method that can directly synthesize B-rep models) train different decoders / generators to independently generate surfaces, curves, and vertices. However, these methods fail to explicitly encode topological relationships, resulting in topological inconsistencies in the generated B-rep models.

[0052] 2. Methods based on multimodal input: For example, CAD-MLLM can generate B-rep models based on multiple inputs such as point clouds, images, and text, but its generation process still relies on multiple steps and there is no unified representation space.

[0053] 3. Graph neural network-based methods: such as BRepNet and AutoMate. Although they can capture complex topological relationships, they still rely on multi-stage processing when generating B-rep models, resulting in high computational overhead.

[0054] Based on the above existing B-rep model generation method, the shortcomings of the existing technology are summarized as follows:

[0055] 1. Lack of topological consistency: Existing methods usually generate different geometric elements of the B-rep model in steps, lacking overall topological constraints, which makes the generated model prone to problems such as curve breakage and vertex loss.

[0056] 2. High proportion of invalid B-rep: Since the topological information is not effectively constrained during the training process, the generated B-rep model has low efficiency. For example, the efficiency of BRepGen is only about 50%.

[0057] 3. High computational overhead: Existing methods rely on multi-stage training, such as training the surface generator and curve generator separately. The training process is complex and the inference time is long.

[0058] 4. Insufficient support for multimodal input: Existing B-rep model generation methods are usually limited to point cloud input and lack compatibility with multiple inputs such as text, pictures, sketches, etc.

[0059] Based on the defects of the existing technology, this embodiment provides a method for generating a boundary representation model based on a structured latent space. The method can be applied to a terminal, which can be an intelligent terminal product such as a computer, a smart TV, or a mobile phone.

[0060] The boundary representation model generation method based on structured latent space proposed in this embodiment is a novel representation method for learning and generating computer-aided design (CAD) models, which is presented in the form of a boundary representation (B-rep) model. This boundary representation method unifies the continuous geometric properties of B-rep primitives (such as surfaces and curves) and their discrete topological relationships in a structured latent space. This method is based on a simple observation: the topological connection between two surfaces is essentially closely related to the geometry of their intersection. This prior knowledge can reformulate the topological learning of B-rep as a geometric reconstruction problem in Euclidean space.

[0061] Specifically, this embodiment eliminates curves, vertices, and all topological connections from the structured latent space and uses a neural intersection network to learn to identify and extract curve geometry from a pair of surface primitives. Therefore, the overall structured latent space of this embodiment is defined only for surfaces, but it can fully encode the entire B-rep model, including the surface geometry, curves, vertices, and their topological relationships.

[0062] The compact and integrated structured latent space proposed in this embodiment enables the design of the first diffusion model-based B-rep generator that can accept a variety of input types, including point clouds, single- and multi-view images, 2D sketches, and text descriptions. The method based on this embodiment significantly reduces the ambiguity, redundancy, and inconsistency problems encountered in the B-rep generation process, and reduces the training complexity of the previous multi-step B-rep learning pipeline. At the same time, it far exceeds the existing state-of-the-art methods in terms of B-rep generation efficiency, achieving an efficiency of 82%.

[0063] Specifically, if Figure 1 As shown in , the boundary representation model generation method based on structured latent space of this embodiment includes the following steps:

[0064] Step S100: Encode continuous geometric information and corresponding discrete topological information from different primitive types to form a unified and expressive latent vector, wherein the primitive types include surfaces, curves, and points.

[0065] To obtain the structured latent vector z from the B-rep model s ,like Figure 2 As shown in , this embodiment encodes the geometric information of different primitive types and the corresponding discrete topological information into the surface latent space based on a preset variational autoencoder (VAE) to obtain the latent vector.

[0066] Specifically, in the geometric encoding process, given m surfaces Sampling points and n curves First, a series of convolution and downsampling layers are applied to the surface and curves superior:

[0067] f s =E gs (S i ),f c =E gc (C i )

[0068] in and is the geometric feature vector of surfaces and curves. Note that the feature dimension of surfaces is designed to be 32 = 2 × 2 × 8, where 8 represents the feature dimension. Because the permutation of the max pooling operation remains unchanged, the orientation information of the primitives is lost. Therefore, this embodiment retains a spatial resolution of 2 to distinguish the orientation of the primitives. This design is also used for all other feature vectors in the network.

[0069] Next, this embodiment trains a graph neural network Based on the topological connection relationship T SC The curve features of are propagated to surface primitives:

[0070] f cs =GNN(f s ,f c ,T SC ),#(2)

[0071] in This example uses a series of self-attention layers to To aggregate and exchange the features in each surface latent space, so as to capture long-term relationships and further strengthen the feature vector. Finally, this embodiment uses a multi-layer perceptron (MLP): The topology-aware surface latent space is mapped to the mean and variance of the Gaussian latent space. Finally, the latent vector z can be sampled from the predicted mean and variance. s .

[0072] Step S200: Decode the latent vectors in the structured latent space to recover the relationships between all primitive types and their geometric and topological information.

[0073] A direct approach to supervised training in existing methods is to reconstruct surface primitives And use L1 norm or L2 norm loss between reconstruction and output surface. However, this does not guarantee that the latent vector will encode topological information and other primitives. In order to force the latent vector to encode the structural information of the B-rep model, this embodiment designs a neural intersection model to recover low-level primitives (curves and points) based on the surface latent vector z s The relationship between geometry and topology, such as Figure 3 shown.

[0074] Specifically, the observation that the intersection of two surfaces is always limited to their boundaries is exploited in this embodiment to insert topological and curve information into the latent space. A pre-defined neural intersection module is used to recover the geometric and topological relationships between curves and points based on the surface latent vectors, and to predict whether the two surfaces intersect. The sampled latent vectors and intersection curve features are then determined, and a decoder network consisting of a CNN and upsampling layers is used to reconstruct surface and curve primitives.

[0075] Specifically, this embodiment first uses a series of We use a self-attention layer to perform deep feature exchange on different surface latent vectors sampled from the latent space. For each pair of surface latent vectors, we add a position encoding based on the surface order (whether it is the first or second surface) to ensure that the intersection is subject to the surface order condition. We then apply a series of cross-attention layers To further exchange the features of the two surfaces, the latent vector of the first surface is used as the query vector, and the second surface is used as the key vector and value vector. Finally, this embodiment uses a The multi-layer perceptron layer maps the exchanged features to the corresponding curve features. This embodiment also trains a A binary classifier is used to determine whether two surfaces intersect.

[0076] Decoding and loss function: given the latent vector z after sampling s and intersection curve feature z c , this embodiment uses a decoder network containing CNN and upsampling layers to reconstruct surface and curve primitives

[0077]

[0078] in and Corresponding reconstructed surface and curve primitives.

[0079] The loss function of this example consists of three parts:

[0080] L1 reconstruction loss of a surface and curve sampling point

[0081] Binary cross entropy loss for an intersection classifier

[0082] A KL regularization loss that forces the latent vector to follow a standard Gaussian distribution

[0083]

[0084] Where w1=1, w2=1e -1 ,w3=1e -6 is the weight of each loss value, μ i is the noise mean, indicating the average level of noise, σ i is the noise standard deviation, which indicates the discrete degree or fluctuation amplitude of the noise. In this embodiment, the learning rate is 1e -4The network is trained using the Adam optimizer until the validation loss converges. Because the number of intersecting sample pairs is far less than the total number of possible sample pair combinations, a balanced sampling strategy is employed to ensure an equal number of positive and negative sample pairs during training. During the inference phase, this embodiment uses an intersection classifier to predict the intersection of all sample pairs.

[0085] During the experiment, the embodiment found that the half-curve structure plays a key role in network robustness training. Although the 16×3 sampling points evenly distributed on the curve are directional in themselves, if the L1 loss function is directly applied to the reconstructed curve and the input curve without considering the half-curve structure, the network will learn a curve with a fixed directionality (the direction is randomly assigned during the data preparation stage). Although the chamfer distance can be used instead of the L1 loss function, this method will significantly increase the computational overhead. More importantly, the directionality of the curve is crucial for forming a closed curve loop for accurately trimming the surface. To this end, this embodiment trains the network based on the order of the surface pairs to predict a directional curve (half-curve). The direction of the true value (GT) of the half-curve is defined based on the first surface in the surface pair. Specifically, when used as the outer boundary of the surface, the half-curve should form a counterclockwise closed loop, and when used as the inner hole boundary, it should form a clockwise loop. If the order of the surface pairs is swapped, the predicted half-curve direction also needs to be reversed accordingly.

[0086] Step S300: training a latent space diffusion model based on the latent vector, and generating a boundary representation model based on the latent space diffusion model according to different input conditions.

[0087] This embodiment pads the latent vector to a fixed length by randomly repeating it until it reaches a predefined maximum number of surface primitives. Then, using the conditional vector as a constraint, the random noise is denoised and mapped to a target latent vector, where the conditional vector is a feature vector generated based on the input condition. Finally, the latent space diffusion model is trained using a denoising diffusion probability model linear scheduler, with the loss function set as the L2 loss between the target latent vector and the true value latent vector.

[0088] Based on the learned latent vector expression This embodiment trains a latent space diffusion model (LDM) to achieve B-rep model generation based on multiple input conditions, such as Figure 3 As shown, this embodiment first repeats the latent vector z randomly s , fill it to a fixed length Until the predefined maximum number of surface primitives M is reached. Then this embodiment adopts a standard transformer architecture With the 256-dimensional conditional input c as the constraint, the random noise Denoising map is the target latent vector

[0089]

[0090] Among them, the conditional vector It is a 256-dimensional feature vector generated based on the input conditions. When there is no input condition, the condition vector is set to zero. If the input condition is a single view, multiple views, or a sketch, a pre-trained DINOv2 model is used to extract a 1024-dimensional feature vector for each input image. When there are multiple input images, the viewpoint pose information is embedded by adding position encoding, and the feature vectors of all images are mean-fused. The MLP layer is used to map the feature vector to a 256-dimensional condition vector. If the input condition is a sparse point cloud or a dense point cloud, a PointNet network is used to extract a 1024-dimensional feature vector for the input point cloud data, which is then mapped to a 256-dimensional condition vector through the MLP layer.

[0091] This example uses the denoising diffusion probability model (DDPM) linear scheduler to train the latent diffusion model (LDM). The configuration parameters include: beta parameter range [0.0001, 0.02], diffusion step size 1000 steps, and fixed learning rate 1e 4 The Adam optimizer of . The loss function is set to predict the latent vector The L2 loss between the hidden vector z and the true value.

[0092] When the denoised latent vector is obtained After that, this embodiment inputs it into the intersection module and the decoder network to finally generate the B-rep model (S, C, T SC ). Similar to the workflow of prior methods such as SolidGe and BRepGen, the output is converted into a watertight CAD model via OpenCascade. For each surface and curve, a C2-smooth B-spline primitive is first fitted to the recovered sample points. The common endpoints of curves on the same surface are then detected and connected as coils. It is important to note that each surface can have multiple coils, with the largest coil defining the outer boundary of the surface and the remaining coils defining the inner hole of the surface. Directed coils are used to crop and recover the surfaces. Finally, all surfaces are stitched together into a watertight B-rep model.

[0093] Even when the conditional input is fixed, the generative model has randomness and diversity when outputting the boundary representation (B-rep) model. To improve the quality of the generative model, this embodiment introduces a test-time enhancement strategy. For a given point cloud data, the generative model is run multiple times, each time with different random noises to generate multiple candidate B-rep models. The chamfer distance between each candidate model and the input condition is calculated, and the model with the smallest distance is selected as the final output. Thanks to the parallelization of the denoising process and post-processing steps, the test-time enhancement strategy does not significantly increase the overall inference time.

[0094] The present invention comprehensively evaluates the B-rep models generated by the method of this embodiment from the aspects of legitimacy, light field distance, coverage, maximum average difference, JS divergence, cyclomatic complexity, and mean curvature. The evaluation results of the B-rep models generated by the method of the present invention when there are no input conditions are shown in Table 1 below. Quantitative analysis and comparison of the results generated unconditionally by DeepCAD, BRepGen, and the method of the present invention on the DeepCAD dataset. All values are the average results of 10 runs, with 3000 models sampled from each method in each run and compared with 1000 real models in the test set.

[0095] Table 1

[0096]

[0097]

[0098] It can be seen that when there is no input condition, the B-rep model generated by the method of this embodiment is superior to the B-rep models generated by DeepCAD and BRepGen in terms of legitimacy, coverage, maximum average difference, JS divergence, cyclomatic complexity and average curvature.

[0099] When input conditions are present, the B-rep model generated using the method of the present invention is evaluated across the following modalities: point cloud conditions, text conditions, and image conditions (multi-view / single-view / sketch). When the input condition is point cloud conditions, the present invention uses unconditional generative novelty analysis based on light field distance to measure the similarity between the generated sample and the DeepCAD training set. As shown in Table 2, a lower light field distance indicates a more similarity (lower novelty) between the generated sample and the training set.

[0100] Table 2

[0101]

[0102] The quantitative results of the B-rep model generated under point cloud conditions on the DeepCAD dataset are compared in Table 3.

[0103] Table 3

[0104]

[0105] This demonstrates the ability to generate quantitative results under point cloud conditions with imperfect input data. The method of this embodiment maintains competitive performance even in the face of various types of noise and data defects. Compared to other baseline methods, the proposed holistic implicit representation promotes a more consistent relationship between topology and geometry, resulting in more reasonable generation results.

[0106] Table 4 provides a detailed quantitative comparison of conditional point cloud generation results for the DeepCAD dataset. This table compares the number of vertices, curves, and faces in the generated samples relative to the ground truth, along with metrics for reconstruction accuracy and completeness of single-primitive geometry. We also report the precision and recall of the generated samples in both the geometric and topological dimensions.

[0107] Table 4

[0108]

[0109] Qualitative results are generated based on the scanned point cloud conditions obtained by a structured light scanner. Such point clouds often suffer from 1) noise, 2) localized missing points, and 3) misalignment due to limitations of the scanning equipment and environment. Despite this, the method of this embodiment can still generate multiple reasonable candidate generation solutions even when the input point cloud is severely noisy. Table 5 shows the quantitative results of the B-rep model generated when the input condition is an image.

[0110] Table 5

[0111]

[0112] In summary, the efficiency of generating B-rep models by the method of the present invention reaches 82% to 84%, while the efficiency of the existing BRepGen method is only about 50%. The overall latent space proposed by the present invention clearly encodes the topological relationship, avoiding the topological inconsistency problem caused by unclear topological information generated in the prior art. In addition, the present invention avoids the error accumulation problem that occurs in the BRepGen method during the step-by-step generation (surface → curve → vertex) process; the present invention only requires a single training to generate a complete B-rep model, which reduces computational costs and improves efficiency. The method of the present invention is compatible with multimodal inputs including point clouds, images, text, sketches, etc., while BRepGen mainly focuses on a single input modality and is difficult to adapt flexibly.

[0113] Based on the above embodiments, the present invention further provides a system for generating a boundary representation model based on a structured latent space, wherein the system is used to implement the steps of the method for generating a boundary representation model based on a structured latent space in any one of the above method embodiments, such as Figure 4 As shown in , the system includes: a latent vector encoding module 10, a latent vector decoding module 20, and a boundary representation model generation module 30. Specifically, the latent vector encoding module 10 is used to encode continuous geometric information and corresponding discrete topological information from different primitive types to form a unified and expressive latent vector. The primitive types include surfaces, curves, and points. The latent vector decoding module 20 is used to decode the latent vectors in the structured latent space to restore the relationship between all primitive types and their geometric information and topological information. The boundary representation model generation module 30 is used to train a latent space diffusion model based on the latent vectors, and generate a boundary representation model based on the latent space diffusion model according to different input conditions.

[0114] The working principles of each module in the boundary representation model generation system based on structured latent space in this embodiment are the same as the principles of each step in the above method embodiment, and will not be repeated here.

[0115] Each module in the above-mentioned structured latent space-based boundary representation model generation system can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a terminal in the form of hardware, or can be stored in a memory in the terminal in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0116] Based on the above embodiment, the present invention further provides a terminal, the principle block diagram of the terminal can be as follows: Figure 5 The terminal may include one or more processors 100 ( Figure 5 Only one is shown in the figure), a memory 101, and a computer program 102 stored in the memory 101 and executable on one or more processors 100. For example, a boundary representation model generation program based on a structured latent space. When one or more processors 100 execute the computer program 102, each step in the embodiment of the method for generating a boundary representation model based on a structured latent space can be implemented. Alternatively, when one or more processors 100 execute the computer program 102, the functions of each module / unit in the embodiment of the system for generating a boundary representation model based on a structured latent space can be implemented, which is not limited here.

[0117] In one embodiment, the processor 100 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0118] In one embodiment, the memory 101 may be an internal storage unit of an electronic device, such as a hard disk or memory of the electronic device. The memory 101 may also be an external storage device of the electronic device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. Furthermore, the memory 101 may include both an internal storage unit of the electronic device and an external storage device. The memory 101 is used to store computer programs and other programs and data required by the terminal. The memory 101 may also be used to temporarily store data that has been output or is about to be output.

[0119] Those skilled in the art will understand that Figure 5 The principle block diagram shown in the figure is only a block diagram of a partial structure related to the solution of the present invention, and does not constitute a limitation on the terminal to which the solution of the present invention is applied. The specific terminal may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0120] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, operating database or other media used in the embodiments provided by the present invention may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct RAMbus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.

[0121] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for generating a boundary representation model based on a structured latent space, characterized in that: The method comprises: Encodes continuous geometric information from different primitive types, including surfaces, curves, and points, and the corresponding discrete topological information into a unified and expressive latent vector. Decode the latent vectors in the structured latent space to recover the relationships between all primitive types and their geometric and topological information; A latent space diffusion model is trained based on the latent vector, and a boundary representation model is generated based on the latent space diffusion model according to different input conditions.

2. The method for generating a boundary representation model based on a structured latent space according to claim 1, characterized in that: The encoding comes from continuous geometric information of different primitive types and the corresponding discrete topological information, forming a unified and expressive latent vector, including: Based on a preset variational autoencoder, the geometric information of different primitive types and the corresponding discrete topological information are encoded into the surface latent space to obtain the latent vector.

3. The method for generating a boundary representation model based on a structured latent space according to claim 2, characterized in that: The preset variational autoencoder is used to encode the geometric information of different primitive types and the corresponding discrete topological information into the surface latent space to obtain the latent vector, including: Determine several sampling points of the surface and several curves; Train a graph neural network to propagate curve features based on topological connectivity to surface primitives; Use a series of self-attention layers to aggregate and exchange features in the latent space of each surface; A multi-layer perceptron is used to map the topologically aware surface latent space into the mean and variance of the Gaussian latent space, and the latent vector is sampled from the predicted mean and variance.

4. The method for generating a boundary representation model based on a structured latent space according to claim 1, wherein: The decoding of latent vectors in the structured latent space to recover the relationship between all primitive types and their geometric and topological information includes: Based on the preset neural intersection module, it recovers the relationship between the geometric information and topological information of curves and points based on the surface latent vector, and predicts whether two surfaces intersect; The sampled latent vectors and intersection curve features are determined, and a decoder network consisting of CNN and upsampling layers is used to reconstruct surface primitives and curve primitives.

5. The method for generating a boundary representation model based on a structured latent space according to claim 4, characterized in that: The method of restoring the relationship between the geometric information and topological information of curves and points based on surface latent vectors based on a preset neural intersection module and predicting whether two surfaces intersect includes: Use a series of self-attention layers to perform deep feature exchange on different surface latent vectors sampled from the structured latent space; A series of cross-attention layers are used to exchange the features of two surfaces, where the latent vector of the first surface is used as the query vector and the latent vector of the second surface is used as the key vector and value vector; A multi-layer perceptron layer is used to map the swapped features to the corresponding curve features, and a binary classifier is trained to determine whether the two surfaces intersect.

6. The method for generating a boundary representation model based on a structured latent space according to claim 1, wherein: The method of training a latent space diffusion model based on latent vectors and generating a boundary representation model based on the latent space diffusion model according to different input conditions includes: Pad the latent vector to a fixed length by randomly repeating it until a predefined maximum number of surface primitives is reached; Using the conditional vector as a constraint, the random noise is denoised and mapped into a target latent vector, wherein the conditional vector is a feature vector generated according to the input condition; The latent space diffusion model is trained using a denoising diffusion probability model linear scheduler, and the loss function is set to the L2 loss between the target latent vector and the true value latent vector.

7. The method for generating a boundary representation model based on a structured latent space according to claim 6, characterized in that: The condition vector is used as a constraint, including: If there is no input condition, the condition vector is set to zero; If the input condition is single view, multi-view, or sketch, for each input image, a 1024-dimensional feature vector is extracted using the pre-trained DINOv2 model. When there are multiple input images, the view pose information is embedded by adding position encoding, and the feature vectors of all images are mean-fused. The MLP layer is used to map the feature vectors to a 256-dimensional conditional vector. If the input condition is a sparse point cloud or a dense point cloud, a 1024-dimensional feature vector is extracted from the input point cloud data using the PointNet network, and then mapped to a 256-dimensional conditional vector through the MLP layer.

8. A boundary representation model generation system based on structured latent space, characterized in that: The system is used to implement the steps of the method for generating a boundary representation model based on a structured latent space according to any one of claims 1 to 7, and the system includes: A latent vector encoding module, which encodes continuous geometric information from different primitive types, including surfaces, curves, and points, and the corresponding discrete topological information, to form a unified and expressive latent vector; The latent vector decoding module is used to decode the latent vectors in the structured latent space to recover the relationship between all primitive types and their geometric and topological information; The boundary representation model generation module is used to train a latent space diffusion model based on latent vectors, and generate a boundary representation model based on the latent space diffusion model according to different input conditions.

9. A terminal, characterized in that: The terminal includes a memory, a processor, and a boundary representation model generation program based on a structured latent space stored in the memory and runnable on the processor. When the processor executes the boundary representation model generation program based on a structured latent space, the steps of the boundary representation model generation method based on a structured latent space as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a boundary representation model generation program based on a structured latent space. When the boundary representation model generation program based on a structured latent space is executed by a processor, the steps of the boundary representation model generation method based on a structured latent space as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Three-dimensional model representation method and system for expressing geometric details and complex topology

    CN110889893A

  • Implicit nerve characterization method based on geometric primitives

    CN118864760A

  • Systems and methods of hierarchical implicit representation in octree for 3D modeling

    US20230005217A1

Cited By

  • Boundary representation generation method and system based on graph diffusion and storage medium

    CN120724508A