Computer-aided design sequence and knowledge inference from product images
By using machine learning and AI to generate CAD sequences from image inputs, the system addresses the limitations of current reverse engineering technologies, enabling efficient and accessible CAD model reconstruction.
Patent Information
- Application Number
- PCT/US2024/054373
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-09
- Filing Date
- 2024-11-04
- Publication Date
- 2025-05-08
AI Technical Summary
Current reverse engineering technologies focus on reconstructing 3D CAD models from point clouds, which is labor-intensive and lacks the ability to capture the sequence and intent of the designer's process, as well as being inaccessible to users without specialized expertise.
The development of an image-to-CAD-sequence system and method that employs machine learning and artificial intelligence to predict and generate a sequence of CAD operations based on a single image input, encoding both the sequence information and parameters into a latent space, allowing for the creation of 3D CAD models.
This approach enables the reconstruction of CAD models in an automated manner, making the design process more inclusive and accessible, allowing both experienced and novice designers to contribute, and speeding up the design workflow.
Smart Images

Figure US2024054373_08052025_PF_FP_ABST
Abstract
Description
COMPUTER-AIDED DESIGN SEQUENCE AND KNOWLEDGE INFERENCEFROM PRODUCT IMAGESCross-Reference to Related Applications
[0001] This application claims priority to and the benefit of U.S. Provisional Application No. 63 / 595,984, titled “Computer-Aided Design Sequence and Knowledge Inference From Product Images,” filed on November 3, 2023, and U.S. Provisional Application No. 63 / 551,990, titled “Computer-Aided Design Sequence and Knowledge Inference From Product Images,” filed on February 9, 2024, the contents of which are hereby incorporated by reference herein in their entireties.Statement Regarding Federally Funded Research
[0002] This invention was made with government support under Grant number DUE2207408 awarded by the National Science Foundation. The government has certain rights in the invention.Background
[0003] Computer aid in drafting (CAD) software, computer aid in manufacturing (CAM) software, computer aid in engineering (CAE) are critical tools used by designers and engineers in the design, analysis, and engineering of mechanical parts and systems. Often time, such models and files are shared by the designers and engineers to other organizations, e.g., manufacturers. Issues with outdated documentation and a lack of digital records can arise. In addition, changes to designs can be made at the manufacturer site without updates being communicated back to the original designers.
[0004] Companies also engage in Reverse engineering (RE) to reconstruct CAD models. Integrating RE with CAD systems allows designers to leverage the advantages of existing products and improve upon the original design. Current RE processes typically have two major limitations. First, they often concentrate on reconstructing three-dimensional (3D) models rather than parametric lists that capture the sequence and intent of a designer in creating those models, leading to the inability to review or access the historical construction process and the associated design knowledge behind the CAD model. Second, they are often primarily performed manually, which is labor-intensive and time-consuming.
[0005] 3D scanners can be used to scan a physical part into a 3D point cloud that is then converted into a CAD model. 3D scanners are specialized equipment that require specialized expertise.
[0006] There is a benefit to generating CAD, CAE, CAM models in automated manner.Summary
[0007] An exemplary image-to-CAD-sequence system and method are disclosed that employ machine leaming / artificial intelligence (Al), used interchangeably herein, trained via vectorized map or matrix of CAD operation and parameters, to predict / generate a sequence of computer-aided design (CAD) operations (“CAD sequence” in short), also referred to as a parametric list of a model, based on a single image input (e.g., rasterized image). Indeed, the exemplary system and method can operate on a rasterized image comprising pixels as input to encode, via a trained Al model, both the sequence information of CAD operations and their corresponding parameters into a latent space, capable of capturing high-level information given the training data of CAD sequences.
[0008] The term “CAD,” as used herein, in the context of models, objects, and operations, includes models, objects, and operations for computer-aided design (CAD) as well as computer-aided manufacturing (CAM) and computer-aided engineering (CAE).
[0009] Rasterized images are analyzed and employed to generate a vectorized map or matrix that includes a set of CAD operations and corresponding parameters. The vectorized map or matrix is employed with the rasterized images as training data for the Al system, which then can generate a vectorized map or matrix from a new rasterized image during runtime operation. The vectorized map or matrix is translatable to a sequence of predefined operations and associated parameters, e.g., a parametric list.
[0010] The exemplary system and method can make the CAD model reconstruction process more inclusive and easily understandable to a broader range of users, to aid both experienced and novice designers to actively contribute to the design, thus promoting design collaboration. Moreover, it has the potential to involve customers in the design process, fostering design democratization. It can also speed up the design workflow.
[0011] Potential benefits of the exemplary system and method include at least features. First, different from the current technologies that focus primarily on generating 3D models, the exemplary system and method can produce CAD sequences that can ultimately lead to the creation of 3D models. Additionally, CAD sequences offer a chronologicalprocedure that illustrates the step-by-step construction of 3D models that embed design knowledge, thus providing valuable design insights and educational benefits. Second, the systems and methods described herein eliminate the requirement for complex tools to obtain point clouds as input, which is typically required by current technologies for reverse engineering. Instead, it only relies on the images of the 3D objects, which can be more easily acquired compared to point clouds. Using CAD sequences offers several advantages compared to 3D CAD models, which grant greater flexibility in modifying individual steps and enhances the comprehension of the historical process and design knowledge / experience behind constructing the CAD model.
[0012] In some embodiments, a designer can provide a sketch or an image of a design to which a CAD, CAM, or CAE model with parametric lists can be generated for subsequent modification. The generated vectorized map or matrix provides a parametric list of CAD- software-specific operations that are compatible and native to existing CAD software.
[0013] In some embodiments, the exemplary method employs a target-embedding variational autoencoder (TEVAE). Other generative Al-based architecture may be employed.
[0014] In an aspect, a method is disclosed for generating a parametric list for a computer-aided design (CAD) application. The method includes: providing at least one image of an object; generating, using a trained machine learning model and based on the image, a vectorized object corresponding with the object; generating, using the trained machine learning model, the parametric list based on the vectorized object, wherein the parametric list includes a sequence of predefined operations and associated parameters to construct a corresponding CAD object, and wherein the parametric list is used, and is modifiable, in a CAD software or CAD workflow to construct a three-dimensional (3D) model of the CAD object.
[0015] In some implementations, the trained machine learning model is trained using a training data set of images and corresponding vectorized objects and / or sequences of CAD operations.
[0016] In some implementations, the machine learning model is trained in a first training stage and a second training stage.
[0017] In some implementations, the first training stage includes unsupervised learning, wherein the second training stage includes supervised learning.
[0018] In some implementations, the trained machine learning model includes a target-embedding variational autoencoder (TEVAE) architecture.
[0019] In some implementations, the TEVAE architecture includes a first encoderdecoder sub-model (e.g., variational autoencoder (VAE) with transformer-based blocks) that is configured to encode sequence information of the predefined operations and their corresponding parameters into a latent space.
[0020] In some implementations, the TEVAE architecture includes a second encoder sub-model that is configured to regress the latent space using images (i.e., input feature objects).
[0021] In some implementations, the parametric list includes between 6 and 30 discrete and / or continuous parameters.
[0022] In some implementations, each parameter corresponds with a user-specified predefined value range (e.g., specified by a developer / researcher). User-specified inputs includes an input that is received through an interface, e.g., via a customization screen.
[0023] In some implementations, the parameters include one or more of a operation type, an identifier of a sketch plane, at least one end coordinate point x, at least one end coordinate point y, a sweep or resolve angle, a radius, an identifier of the profile, an identifier of reference, an extrusion distance, Boolean operations, a scale factor, a mode, a Rho value and a width.
[0024] In some implementations, the CAD object is employed in computer-aided engineering (CAE) analysis or computer-aided manufacturing (CAM) analysis.
[0025] In some implementations, the CAD object is employed in a generative artificial intelligence model.
[0026] In some implementations, a system is provided. The system can include: a processor; and a memory having instructions thereon, wherein the instructions when executed by the processor, cause the processor to: receive at least one image of an object; generate, using a trained machine learning model and based on the at least one image, a vectorized object corresponding with the object; generate, using the trained machine learning model, a parametric list based on the vectorized object, wherein the parametric list includes a sequence of predefined operations to construct a corresponding CAD object, and wherein the parametric list is used, and is modifiable, in a CAD software or CAD workflow to construct a 3D model of the CAD object.
[0027] In some implementations, the instructions when executed by the processor, cause the processor to: train the machine learning model using a training data set of images and corresponding vectorized objects and / or sequences of CAD operations.
[0028] In some implementations, the instructions when executed by the processor, cause the processor to: train the machine learning model in a first training stage and a second training stage.
[0029] In some implementations, the first training stage includes unsupervised learning, and wherein the second training stage includes supervised learning.
[0030] In some implementations, the trained machine learning model includes a target-embedding variational autoencoder (TEVAE) architecture.
[0031] In some implementations, the TEVAE architecture includes a first encoderdecoder sub-model or variational autoencoder (VAE) with transformer-based blocks that is configured to encode sequence information of the predefined operations and their corresponding parameters into a latent space.
[0032] In some implementations, the TEVAE architecture includes a second encoder sub-model that is configured to regress the latent space using images.
[0033] In some implementations, the techniques described herein relate to a non- transitory computer readable medium including a memory having instructions stored thereon, which when executed by a processor, cause the processor to: receive at least one image of an object; generate, using a trained machine learning model and based on the at least one image, a vectorized object corresponding with the object; generate, using the trained machine learning model, a parametric list based on the vectorized object, wherein the parametric list includes a sequence of predefined operations to construct a corresponding CAD object, and wherein the parametric list is used, and is modifiable, in a CAD software or CAD workflow to construct a 3D model of the CAD object.
[0034] Additional advantages of the disclosed systems and methods will be set forth in part in the description that follows and in part will be obvious from the description.
[0035] The advantages of the disclosed compositions and methods will be realized and attained by means of the elements and combinations particularly pointed out in the appended claims. It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and is not restrictive of the disclosed compositions and methods, as claimed.
[0036] The details of one or more embodiments of the present disclosure are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the present disclosure will be apparent from the description and drawings and from the claims.
[0037] Throughout the description and claims of this specification, the word “comprise” and other forms of the word, such as “comprising” and “comprises,” means including but not limited to, and is not intended to exclude, for example, other additives, components, integers, or steps.Brief Description of the Figures
[0038] The components in the drawings are not necessarily to scale relative to each other. Like reference, numerals designate corresponding parts throughout the several views.
[0039] Figure 1A is an example system in accordance with certain embodiments of the present disclosure.
[0040] Figure IB is an example system in accordance with certain embodiments of the present disclosure.
[0041] Figure 2 is flowchart diagram illustrating a method in accordance with certain embodiments of the present disclosure.
[0042] Figure 3 A is flowchart depicting a method in accordance with certain embodiments of the present disclosure.
[0043] Figure 3B shows the overview of an example network and two-stage training process in accordance with certain implementations described herein.
[0044] Figure 3C shows an example method for recognizing a parametric list and translating it to a 3D CAD model in accordance with certain implementations described herein.
[0045] Figure 3D shows an example network in accordance with certain embodiments described herein.
[0046] Figure 3E shows an automatic data synthesis pipeline in accordance with certain embodiments described herein.
[0047] Figure 4A shows outputs of an example data synthesis pipeline including a set of template shapes (TS) that were developed.
[0048] Figure 4B shows two implementations of the Image2CADSeq network.
[0049] Figures 5 A-5D are graphs depicting the accuracy of operation parameters in relation to parameter tolerance.
[0050] Figures 5E-5G are graphs providing a comparison of the network's performance evaluated under two different architectures and two datasets.
[0051] Figure 5H and Figure 51 are graphs illustrating the relationship between parameter accuracy and tolerance.
[0052] Figure 5J, Figure 5K, and Figure 5L are each a series of graphs depicting experimental results.
[0053] Figure 5M shows the summary of the parsing rate, intersection over union (loU), and mean squared error (MSE) outlined in Figure 4B of the three cases.
[0054] Figure 6 is an example computing device.Detailed Specification
[0055] To facilitate an understanding of the principles and features of various embodiments of the present disclosure, they are explained hereinafter with reference to their implementation in illustrative embodiments.
[0056] In an implementation, a data-driven approach is provided that can generate CAD, CAM, and / or CAE sequences based on a single image. The exemplary system and method can eliminate the need for complex tools and specialized equipment to obtain point clouds, as is commonly required in existing RE technologies.
[0057] Embodiments of the present disclosure relate to the domain of 2D-to-3D generation, led by advancements in computer graphics and computer vision. For an in-depth understanding of the current developments and challenges in this area, readers are directed to a comprehensive review
[0011] , The use of 3D representations, such as voxels, point clouds, and meshes, has been prevalent.
[0058] Despite their widespread use, these representations often focused on generating visually appealing objects without necessarily considering their engineering aspects, such as dimensions, engineering performance, and compatibility with engineering software
[0012] , This mismatch hinders the seamless integration with downstream applications, such as editing and engineering analysis of synthesized 3D shapes, underlining the necessity for adopting CAD-specific data formats.
[0059] Constructive solid geometry (CSG) is one of the fundamental methods for creating CAD models. It applies Boolean operations (e.g., union, intersection, and subtraction) to basic geometric shapes (primitives), such as cuboids, spheres, and cylinders. CSG is known for its lightweight structure, which allows for easy modifications by altering the parameters of these primitives and their spatial transformations.
[0060] There have been various studies on the deep learning methods of CAD model generation using CSG [13-16], Despite its merits, CSG lacks the versatility used in contemporary CAD tools that utilize parametric modeling.
[0061] Parametric CAD models start as 2D sketches comprising geometric primitives (e.g., line segments and arcs) with explicit constraints, such as coincidence and perpendicularity, establishing the foundation for 3D construction operations (e.g., extrusionand revolution). The relevant deep learning methods of parametric CAD models include engineering sketch generation, 3D CAD model generation, and reconstruction. There has been a series of studies recently [3, 17-20] dedicated to the generation of CAD sketches through the application of deep learning approaches. The emphasis in these works is on generating 2D layouts rather than dealing with the generation of 3D components. Embodiments of the present disclosure focus on introducing deep learning methods for 3D CAD model generation and reconstruction.
[0062] Deep Learning of Parametric 3D CAD Models. Boundary representation (B- rep) format is the standard format for representing 3D shapes in CAD, which defines objects based on their boundary surfaces, edges, and vertices connected through specific topology. Numerous learning-based approaches have emerged for the generation of parametric curves
[0021] and surfaces
[0022] , In addition to curve and surface generation, Smirnov et al.
[0023] introduced a generative model for creating topology that combines parametric curves and surfaces to create solid models, depending on predefined topological templates. Various methods have been proposed to enable the direct generation of B-rep models with arbitrary topology [24-26], The focus of this disclosure lies in generating CAD sequences that can be translated into B-rep models using a solid modeling kernel, such as Fusion 360 CAD software.
[0063] Significant progress has been made in the generation of CAD sequences, particularly through leveraging Sketch-and-Extrude modeling operations for the reconstruction of 3D models. Recently, there have been methods [2, 27] for generative models specifically designed for the unconditional generation of CAD sequences. These models aim to autonomously create CAD sequences without relying on specific conditions or inputs. Specifically, Wu et al. [2] presented the first generative model, DeepCAD, that learns from sequences of CAD modeling operations to produce editable CAD designs. By drawing an analogy between CAD operations and natural language, the authors propose to utilize a Transformer
[0028] architecture aiming to leverage the capabilities of Transformer models in understanding and generating sequences, adapting them to the context of CAD design operations.
[0064] Generative models indeed serve as valuable tools for randomly generating a multitude of designs, offering inspiration and exploration of diverse possibilities. However, these models lack the capability to directly incorporate designers’ intent into the generation process. Consequently, the designs generated may deviate from the expectations or specific requirements of designers. This discrepancy highlights the need for mechanisms that allowdesigners to guide or influence the output, ensuring that the generated designs align more closely with their intent and preferences. To that end, several methods have been introduced to allow the CAD sequence generation given the target of B-rep models [29, 30], voxels [8,9], point clouds [6,7], and sketches [31,32], Particularly, Fusion 360 Gym
[0029] was developed to reconstruct a CAD model given a B-Rep model, utilizing a face-extrusion technique that relies on existing planar faces within the BRep model. However, despite showcasing the potential for CAD sequence generation, the face-extrusion method differs significantly from the more natural sketch-extrusion method commonly used by human designers. Moreover, this technique is ineffective when confronted with a lack of available planar or profile data.
[0065] This work aims to fill a research gap in the existing literature by focusing on the task of generating CAD sequences from images. To that end, the Transformer-based autoencoder initially introduced in DeepCAD [2] was expanded, transforming it into a TEVAE architecture
[0010] , The application of Fusion 360 Gym
[0029] serves as a case study, utilizing its domain-specific language tailored for CAD sequences involving Sketch-and- Extrude operations.
[0066] Technical Background. An autoencoder (AE) is a type of neural network that aims to learn a compressed representation of input data
[0033] , It consists of two main parts: an encoder and a decoder. The encoder compresses the input into a lower-dimensional latent space, while the decoder reconstructs the input data from this compressed representation. The goal is to minimize the difference between the original input and its reconstruction, leading to efficient data encoding. Variational Autoencoders (VAEs)
[0034] extend traditional AEs by introducing a probabilistic way to the encoding process. This probabilistic approach allows VAEs not only to reconstruct input data but also to generate new data that is similar to the input. Autoencoders have been applied to uncover the useful underlying structures of data which is typically known as representation learning
[0035] ,
[0067] Representation learning is a set of techniques in machine learning that automatically discovers the representations with reduced dimensionality from raw data for downstream tasks, such as regression or classification
[0036] , While VAEs are particularly known as generative models for their effectiveness in generating complex data, they offer a more robust, regularized, and probabilistic approach to representation learning compared to traditional AEs [10, 37],
[0068] Although most research has focused on employing AEs and VAEs in unsupervised or semi-supervised scenarios, it is worth noting that autoencoders alsodemonstrate utility in supervised contexts
[0038] , Specifically, incorporating an auxiliary feature-reconstruction task proves beneficial in enhancing supervised classification problems
[0039] , which are known as feature-embedding autoencoders (FEAs). Furthermore, Jarrett and Schaar
[0038] propose target-embedding autoencoder (TEAs) and demonstrate its effectiveness theoretically and empirically. As implied by their names, TEAs focus on encoding target information into a latent space, while FEAs encode feature information.
[0069] TEAs have been applied to various problems. Girdhar et al.
[0040] introduce a TL-embedding network, comprising a T-network (an autoencoder network for the horizontal bar and an encoder for the vertical bar) during training. Upon training completion of the T- network, it facilitates the derivation of the L-network, enabling the prediction of 3D voxel shapes from input images. A similar network architecture has also been applied for semantic image segmentation tasks [41, 42], Drawing inspiration from these preceding studies, Li et al.
[0010] propose to use a VAE to replace the AE and form a target-embedding variational autoencoder (TEVAE) architecture, demonstrating its effectiveness in predictive and generative tasks for car and mug design examples.
[0070] In this study, the Image2CADSeq network was constructed by comparing both TEA and TEVAE architectures. The objective is to assess their effectiveness in the prediction task of images to CAD sequences. The results revealed a significant superiority of the TEVAE architecture over the TEA architecture in terms of prediction performance as detailed herein.
[0071] Example System
[0072] The exemplary system includes features to process one or more images and generate a parametric list based on a vectorized object. The parametric list can include a sequence of predefined operations and associated parameters that can be used to construct a corresponding CAD object. In various implementations, the parametric list is used, and is modifiable, in a CAD software or CAD workflow to construct a three-dimensional (3D) model of the CAD object. A CAD sequence can refer to a sequence of CAD operations (e.g., sketchQ, extrudeQ). Each CAD operation consists of two parts: CAD operation type (a.k.a., the name of the CAD operation, e.g., extrude) and its associated parameters in the bracket.
[0073] A direct output of the system and method described herein can be a native CAD format, such as F3D. These native CAD formats are much more versatile than neutral CAD formats (e.g., STEP, STL) and are easily convertible to other CAD formats (e.g., the neutral CAD formats, voxel formats, mesh formats (e.g., OBJ), and point cloud formats, and the like). For example, as described below in the experimental results below, Fusion 360 andthe native CAD format, F3D, were used and can be converted to other 3D design format, STEP, OBJ, STL and so on. It should be understood that the systems and methods described herein can be adapted to a variety of CAD ecosystems for different platforms and applications.
[0074] Figure 1A and Figure IB show example systems in accordance with certain implementations described herein. As shown in Figure 1A and Figure IB, the system 100A, 100B can include a computing system 102 that is configured to receive at least one image 101 (e.g., rasterized image) of an object as an input and generate a vectorized object 105 corresponding with the image 101 using a trained machine learning model 104. The computing system 102 can be embodied as a cloud-based server, remote server, or local server. The computing system 102 may be connected to one or more storage area network devices and may provide services (e.g., web-based services) to generate 3D CAD models 109 by one or more web-based CAD applications executing on other computing devices.
[0075] The machine learning model 104 can be a supervised machine learning model, an unsupervised machine learning model, or combinations thereof. The machine learning model 104 can be trained in a plurality of stages (e.g., Stage 1 and Stage 2, as described in more detail with reference to Figures 3B-3D). The machine learning model 104 can be trained using a training data set of images and corresponding vectorized objects and / or sequences of CAD operations. In some implementations, the machine learning model 104 includes a target-embedding variational autoencoder (TEVAE) architecture. For example, as shown, the machine learning model 104 (e.g., TEVAE) can comprise a first encoder-decoder submodel 120 and a second encoder-decoder submodel 121. The first encoder-decoder submodel 120 can be configured to encode sequence information of the predefined operations and their corresponding parameters into a latent space. The second encoder-decoder submodel 121 can be configured to regress the latent space using images as input feature objects.
[0076] The computing system 102 can generate a parametric list 107 based, at least in part, on the generated vectorized object 105. The parametric list 107 can include a sequence of predefined operations and associated parameters to construct a corresponding CAD object. By way of example, the parametric list 107 can include between 6 and 30 discrete and / or continuous parameters. Each parameter can correspond with a user-specified predefined value range (e.g., specified by a developer / researcher). Example parameters can include operation type, an identifier of a sketch plane, at least one end coordinate point x, at least one end coordinate point y, a sweep or resolve angle, a radius, an identifier of profile, an identifierof reference, an extrusion distance, Boolean operations, a scale factor, a Mode, a Rho value and a width.
[0077] As further illustrated, the parametric list 107 can be used and / or is modifiable in a CAD software 108 or CAD analysis to construct a 3D CAD model 109 of the object in the image 101 (e.g., via a deterministic process using a paring algorithm in CAD language). The CAD software 108 can be, without limitation, a CAD platform, manufacturing system (CAM), or engineering system (CAE). With reference to Figure IB, in some implementations, the parametric list is used to construct a 3D CAD model 109 for a generative Al model 118.
[0078] Example Methods
[0079] Figure l is a flowchart diagram illustrating an example method 200 of generating a 3D CAD model in accordance with certain implementations descried herein. The method 200 can be performed using the systems 100 A, 100B described above in relation to Figure 1A and Figure IB.
[0080] At step 210, the method 200 includes providing (e.g., receiving, obtaining) at least one image of an object.
[0081] At step 220, the method 200 includes generating a vectorized object corresponding with the image. In some implementations, the vectorized object is generated using a trained machine learning model.
[0082] At step 230, the method 200 includes generating a parametric list based on the vectorize object. In some implementations, the parametric list is generated using a trained machine learning model.
[0083] At step 240, the method 200 includes using the parametric list in CAD software or a CAD workflow to construct or generate a 3D model of the CAD object.
[0084] At step 250, the method 200 includes employing the CAD object in analysis, such as, CAE analysis, CAM analysis, a generative Al model or system.
[0085] The flowchart depicted in Figure 3A illustrates the proposed systematic method 300A for predicting CAD sequences 305 given images 101, a process also referred to herein as Img2CADSeq. To effectively tackle this challenge, a clear problem definition that divides the task into manageable components is initiated. The objective is to harness the power of deep learning to predict a CAD sequence 305 — a series of CAD operations 305a characterized by specific operation types 305b and their corresponding parameters 305c — from an image 101. The image 101 could be a rendering from a CAD model or a real -world photograph of a 3D object 130. Due to the intricate nature of CAD sequences 305, a CADprogram 108a is employed as a representational tool. A CAD program 108a enables designers to script their designs programmatically in a specialized scripting environment, such as the Fusion 360 Python API, FreeCAD Python API, or CADQuery. CAD programs 108a can convert CAD sequences 305 described in text into script language that is interpretable and executable by computers. To overcome the inherent lack of structured format in the CAD sequence data, the CAD program 108a is then streamlined into a vectorized representation 318a conducive to neural network processing. This vectorized representation 318a can facilitate not only the development of the neural network architecture 104a described herein but also the creation of a data-synthesis pipeline 320 tasked with generating the training data 322 for the neural network 104a. In addition, given the complexity of the Img2CADSeq task, a comprehensive evaluation system 324 was developed that rigorously assesses the disclosed neural network models’ performance, thereby ensuring the reliability and accuracy of the methodology. This holistic evaluation is crucial in refining the model 104a and guiding its evolution to meet the demanding standards of CAD sequence prediction.
[0086] Figure 3B shows the overview of an example Image2CADSeq network 300B and two-stage training process in accordance with certain implementations described herein. The Image2CADSeq network 300B comprises a target-embedding variational autoencoder (TEVAE) architecture. The network employs a variational autoencoder (VAE) 310 to encode the target objects 305 (CAD sequences) into a latent space, and a separate encoder 312 (e.g., convolutional neural network encoder) is used to regress the latent space using the input feature objects (i.e., images). For the VAE 310, a transformer-based VAE [3] can be used, while a typical ResNetl8 [6] can be employed for the encoder 312.
[0087] Two-Stage Training: A two-stage training approach may be employed, e.g., as described in [5], In Stage 1 302, the VAE 310 is independently trained. Subsequently, in Stage 2 304, the trained VAE’s 310 learning parameters are kept fixed, and the Enc( ) component 312 is trained. By minimizing the Euclidean loss between the latent vectors obtained from the VAE 310 and the embedding vectors generated by Enc( ) 312, the Enc( ) module is trained using input images.
[0088] In some implementations, the first training stage 302 is an unsupervised learning stage, and the second training stage 304 is a supervised learning stage. By way of example, Stage 1 302 is performed to train using all CAD sequences 305 to get a sufficient latent space. The trainable parameters of Stage 1 302 can then be modified / fixed and go to Stage 2 training 304. Each CAD sequence 305 now has a corresponding latent vector in the trained latent space; that is, one data point can now represent the training. Then in Stage 2304, an image 101a (e.g., image rendered from the 3D CAD model obtained from its corresponding CAD sequence) is used as input and an Al / machine learning model such as, but not limited to, a convolutional neural network (CNN) encoder 312, is trained to regress the latent vector to, for example, minimize the difference between the output from the CNN encoder 312 and the latent vector. In the second training stage 304, the input feature x is image and label y is the latent vector obtained in the first training stage 302. The latent vector can be or comprise a compact vectorized design representation of its CAD sequence 305.
[0089] Figure 3C shows an example method 300C in which a user-specified rule 320 and a CAD protocol 324 is utilized by a computing device (e.g., computing device 600 described below in connection with Figure 5) to recognize a parametric list 107 and translate it to a 3D CAD model 109. As depicted, a user-specified rule 320 or parsing rule is used to translate a vectorized design representation 318 to a parametric list 107 (e.g., specified by a developer / researcher).
[0090] Neural Network Architecture, Training, and Application. The application of target-embedding representation learning in deep learning, particularly for cross-modal tasks, has shown considerable efficacy [12, 38], Figure 3D shows an example Image2CADSeq network 300D is in accordance with certain embodiments described herein. As shown, the Image2CADSeq network 300D utilizes a target-embedding representation learning approach. It features an encoder-decoder network 314 for Stage 1 (SI) 302, which is geared towards unsupervised learning and enables the efficient encoding of target objects (i.e., matrix feature of CAD programs) within a latent space. An additional encoder 316 is integrated into Stage 2 (S2) 304, focusing on supervised learning to regress the previously learned latent space using feature objects (i.e., images) as input.
[0091] The Image2CADSeq network 300D employs a two-stage training strategy [10,41], In Stage 1 302, the focus is on independent training of the encoder-decoder network 314. The objective is to minimize the reconstruction loss between the actual matrix feature of CAD programs (y) and its reconstructed equivalent (y ' ). Completing this stage involves fixing the learnable parameters and saving the model, thereby capturing a latent space of y. Stage 2 304 shifts the focus to independent training of the S2 encoder 316 by minimizing the discrepancy between the latent vector (z), derived from the learned latent space, and the embedding vector (z’) produced by the S2 encoder 316 using an image (x) as input. Importantly, each image used in this stage is directly associated with its feature matrix (y)from Stage 1 302. This image (x) and its corresponding feature matrix (y) are associated with the same 3D object, and they form one data pair.
[0092] The alignment of the latent vector (z) with the embedding vector (z’) is performed specifically for these data pairs, ensuring that the S2 encoder 316 training is precisely tuned to the corresponding images. This approach ensures a cohesive and targeted learning process. The present disclosure provides a novel data synthesis pipeline to generate training data pairs as described in more detail below After training the Image2CADSeq network 300D, the S2 encoder 316 is integrated with the decoder 314b, creating the application module 306. This module 306 is capable of predicting a feature matrix given an image input. Subsequently, this feature matrix can be translated into a CAD program using the Sim-Gallery DSL. Finally, the CAD program can be parsed into a 3D object utilizing Fusion 360 software. With the proposed design representation of the CAD programs and Fusion 360 software, an automatic data synthesis pipeline is introduced, as illustrated in Figure 3E. This method is tailored to generate training data pairs 342 comprising feature matrices of CAD programs 315 and the corresponding images 101a, essential for training the Image2CADSeq network 300D. The process begins with a list of basic shape templates 340a (e.g., template shape primitives), such as cylinders, employing the Sketch-Extrude paradigm of the Gallery DSL. For these basic shapes 340a, a sequence of operations 340 is established, for example, a series of template operations (e.g., add line, add circle). The corresponding parameter values 340c of these operations are then generated. In some implementations, the corresponding parameter values 340c are based on the range specified in Table 3 below. By integrating these template operations with their respective parameters 340c, a complete CAD program 315 is formulated, which is then translated into 3D CAD models 109 through Fusion 360 software. These models 109 are rendered to obtain their images. Additionally, the CAD programs 215 are vectorized and quantized to derive their vectorized representations 318 (e.g., feature matrices or vectorized object). An image 101a paired with its vectorized representation 318 (e.g., feature matrix), both derived from the same CAD program 315, constitutes a data pair 342 (x, y). The method is exemplified using a cylinder model discussed in relation to Figure 4A.
[0093] Experimental Results
[0094] Studies were conducted to evaluate the exemplary systems and methods described herein. In some studies, a particular domain-specific language (DSL), namely Fusion 360 Gallery (abbreviated as Gallery for conciseness)
[0029] was employed for the CAD program, as a solid case to demonstrate the disclosed methodology. Therefore, before thepresentation of how the vectorized design representation for the CAD sequence data was devised, an introduction to the Gallery DSL is provided below.
[0095] Fusion 360 Gallery Domain-Specific Language.
[0096] Table 1 presents a summary of the core elements (i.e., CAD-related elements) in Gallery DSL, which enables the representation of a 2D / 3D design as a CAD program, and Python is used to implement the CAD operations
[0029] , Gallery DSL now supports two major types of CAD operations: Sketch and Extrude. Each CAD operation is de- composed into two fundamental components: the operation type and its corresponding parameters. These elements are analogously mirrored in the Gallery DSL as function names and their associated parameters. A Sketch operation includes the definition of a Sketch Plane and Curves on it. A Sketch Plane can be created by the add sketch(L) function, where I is a plane identifier that can be specified from the three canonical planes ”XY”, ”XZ”, or ”YZ” or other planar faces (e.g., the side face of a cube) present in the current geometry. The Sketch Plane can then be used as the reference coordinate system in 2D for specifying the coordinates. A sequence of Curves, including Line, Arc, or Circle, can be drawn using add line(N, N, N, N), add arc(N,N,N,N,N), and add circle(N,N,N), respectively, Where A is a real number representing the required parameters for a particular operation. As two numbers are needed to define one point, Line uses four numbers for start and endpoints; Arc needs five numbers: start point, center point, and sweep angle; and Circle is specified with three numbers: two for position and one for radius. Executing a Sketch can result in enclosed regions, termed profiles in CAD language. An Extrude operation can extrude a profile from 2D into 3D by us- ing add extrude^!, N, O), where I is an identifier for the profile, and A is a signed number defining the depth of extruding along the normal direction of the profile. The Boolean operation (0) specifies the behavior of the extruded 3D volume, e.g., add to or subtract from other 3D bodies.
[0097] While Gallery DSL currently does not support certain objects, such as spheres and springs, it covers a vast range of them by using expressive Sketch and Extrude operations with Boolean capability. Therefore, it is a good starting point for supporting learning-based methods
[0029] , In the future, the other CAD operations, such as Revolve, Sweep, and Fillet, can be added to the Gallery DSL to expand its design grammar for a full- fledged CAD tool. In this study, the current status of the Gallery DSL is leveraged by only considering the Sketch and Extrude operations.
[0098] Gallery DSL acts as a simplified interface to the more complex Fusion 360 Python API. In essence, it democratizes access to sophisticated CAD design through a more intuitive Python-based interface, effectively bridging the gap between complex CAD operations and the user’s ability to execute them efficiently. For instance, as demonstrated in Table 2, when creating a cylinder with a base circle radius of 5 and a height of 10, the CAD program using Gallery DSL requires only about two-thirds of the effort needed with the Fusion 360 Python API and does not require sophisticated definitions of various variables. This makes the Gallery DSL code easier to operate, and more user-friendly and accessible. However, Gallery DSL still requires a substantial amount of coding efforts in non-CAD-related elements, beyond the core functions that are listed in Table 1 below. This disclosure contemplates additional operation types in the Gallery DSL such as, but not limited to, Resolve and Sweep.Table 1. Fusion 360 Gallery Domain Specific Language
[0099] To further improve its readability and accessibility, the Gallery DSL is simplified by isolating its key CAD-related functions, as highlighted in Table 2, which is referred to herein as Sim-Gallery DSL in the following for briefness. This simplification reduces the process to just three steps for creating a cylinder: add sketcht / ), add circle(-), and add extrude(-). Additionally, a parsing method was developed in Python to convert this simplified version back to the standard Gallery DSL. This ensures compatibility with Fusion 360 CAD soft- ware, facilitating seamless integration and execution of the CAD programs written in the Sim-Gallery DSL. It can also facilitate the vectorized representation for theCAD sequence data as will be introduced below. In addition, the Sim-Gallery DSL was formed following the concept of parametric modeling. In parametric modeling, a design is a sequence of operations that progressively modify the current geometry of the object, which is well represented in the Sim- Gallery DSL by a series of pure CAD operation functions.Table 2: Comparison of CAD programs for creating a cylinder with a base circle radius of5 and a height of 10 using Fusion 360 Python API, Fusion 360 Gallery DSL, and the Simplified Gallery DSL.
[0100] In one study that was conducted, the Fusion 360 Gallery domain-specific language (referred to as Gallery DSL hereafter for brevity) [7] was leveraged for the CAD sequences. In each Gallery DSL command, there are two components: the operation type and its corresponding parameters. However, the number of parameters varies for different operations, which presents challenges in devising a suitable vectorized design representation for the network. Drawing inspiration from the works [3], [8], a fixed-dimensional vector was employed that utilizes 10 parameters, as outlined in Table 3 below, to encode each operation.
[0101] Design Representation of CAD Programs. A standardized design representation is essential for neural networks to effectively interpret CAD programs. Thus, it becomes crucial to devise an efficient method to represent each CAD operation as well as the entire CAD program. There are three major challenges: (1) Diversity in CAD operations: Different CAD programs comprise varying numbers of operations. (2) Variability in parameters: Different CAD operations involve different numbers of parameters. (3) Type of parameters: Parameters can be either continuous or discrete values. To tackle these challenges, a design representation with a unified data structure is proposed. 10 variables (t, I, x, y, a, r, [7], d, O, s') were identified from the Sim-Gallery DSL, detailed in Table 3. In what follows, the approach to handling these variables is described in more detail.
[0102] (1) t G {0, 1, 2, 3, 4, 5, 6} represents the operation types with 0 -4 representing add sketch, add line, add arc, add circle, and add extrude. The values 5 and6 are used to represent the start SOP') and the end (EOP') of a CAD program, which are not typical CAD operations but are included for the learning process to indicate a complete CAD program.
[0103] In some implementations, sequence lengths can be longer than the above example. For example, a sequence can comprise between 6 and 12 values. In otherimplementations, CAD sequences can be extended by generating a CAD sequence that comprises multiple groups of operations.
[0104] (2) 16 {0, 1, 2} indicates the Sketch Plane using one of the canonical planes:”XY”, ”XZ”, or”YZ”
[0105] (3,4) x and are the coordinates of the endpoint for Line and Arc, while they represent the center point when the operation type is Circle. The start point required by Line and Arc was excluded from the design representation by obtaining it from the precedent curve to make sure all curves are connected one after another, making the vectorized representation more compact. There are two extra considerations for this setting: (i) If one curve has no precedent, its start point is defaulted to the origin (0, 0) when parsing the design representation, (ii) For Arc that requires a center point instead of an endpoint, the coordinates for the center point are calculated based on its start point, endpoint, and sweep angle.
[0106] (5) a represents the sweep angle of an Arc.
[0107] (6) r is the radius of a Circle.
[0108] (7) [7] represents the profile index in the Sketch.
[0109] (8) d represents the signed distance of the depth for Extrude.
[0110] (9) O E {0, 1, 2, 3} is used to indicate the Boolean operations: join, cut, intersect, or add, respectively.[OHl] (10) 5 is an auxiliary factor that can be used to scale a CAD model.
[0112] In addition, to standardize the treatment of both continuous and discrete parameters, inspired by [2, 43], continuous parameters are discretized through quantization. This involves: (a) Confining continuous values to a subset of [-1, 1] (e.g., (0, 1] for radius and [-1, 1] for endpoint x and y. (b) Di- viding each range into 256 equal segments, enabling representation as 8-bit integers (i.e., 0 -255). (c) For the sweep angle (a), it is multiplied by 180 during interpretation, (d) Handling scale factor (5): Although the scale factor can be a non-negative continuous value, it is limited to 256 levels for consistency with other continuous values’ quantization. Consequently, the 10 parameters can encode both the operation type and its associated parameters. From the 10 parameters, a fixeddimensional vector can be formalized as a unified design representation for each CAD operation, and the unused parameters will be filled with values of -1.
[0113] The subsequent consideration involves standardizing CAD programs of varying sizes (the number of CAD operations involved). For example, besides the start and end marks, SOP and EOP, a cylinder can be created in three operations as shown in Table 2, while a triangular prism in five operations (add sketch( add line( ) * 3, and addexlrude( )). To achieve a consistent data structure across all CAD programs, a treatment, called maximum program length is introduced. Then, CAD programs shorter than this maximum length are extended by appending end marks until they reach the predetermined length. In this study, 7 variables were utilized to construct a 7- dimensional vector \t,I,x,y, a, r, d\ for each CAD operation. Additionally, default values were assigned to the other three variables [7], O, and s, setting them as 0, 3, and 10 correspondingly. Different variables can be selected which will influence the complexity of the data structure and thus the complexity of designs. In addition, the maximum length for CAD programs is set to 10. As a result, the design representation of a CAD program will be a matrix, namely, the feature matrix.
[0114] Mathematically, the feature matrix denoted as P, is expressed as P = o^, o^, . . . , o^c e R10x7, where o1G R7is a CAD operation vector, and Ac = 10 is the sequence length of the CAD program. Refer to Figure 4 A for an example of how a cylinder is converted to a feature matrix, including the its parameters as explained earlier. In some embodiments, a concatenation of feature matrices can be used to accommodate multiple groups. In the above example, the resulting data representation, P E IR30*7, is a concatenation of three feature matrices representing that the CAD sequence is composed of three groups of operations.Discrete 0-255 (e.g., 10)Table 3: Representative Parameters t, / , x, y, a, r, [I], d, O, s
[0115] There are different types of parameters: some are continuous values (e.g., coordinates x and ), while others are discrete values (e.g., the Boolean operation O). To handle both types of parameters uniformly, a common approach is to convert the continuous values into discrete values through quantization [3], [8], To achieve this, the range of continuous values is restricted to a subset of [-1, 1] (e.g., (0, 1] for radius). Then, each range is divided into 256 equal segments, allowing us to represent them as 8-bit integers ranging from 0 to 255. Especially, for the sweep angle a, it is multiplied by 180 during interpretation. Notably, the scale factor 5 can be any non-negative continuous value, but for consistency, it is limited to 256 levels.
[0116] Vectorized Design Representation. A unified design representation for each Gallery command can be achieved by utilizing a fixed-dimensional vector. In this study, a 7-dimensional vector [t, I, x, y, a, r, d| was employed (highlighted in Table 3) as part of the data synthesis pipeline, as described in connection with Figure 3C above. When interpreting the 7- dimensional vectorized representation, [7], O, and s are defaulted to values of 0, 3, and 10, respectively. In future work, all parameters will be incorporated for a more comprehensive representation, thereby increasing the complexity of the synthesized data. In some implementations, the identifier (1) can represent or indicate a position of a group of operations in a CAD sequence that comprises multiple groups of operations.
[0117] To ensure consistency across CAD programs with varying operation sequence lengths, there is a need for a fixed total number of operations. For instance, a cylinder can be represented by five operations: [5, 0, 3, 4, 6], while a triangular prism may require seven operations: [5, 0, 1, 1, 1, 4, 6] (recall Table 3). To achieve uniformity, 6 (i.e., EOS) is appended to the end of the sequence until it reaches a fixed length (Ni). In the instant study, Ni is set to 10, ensuring it is greater than the longest sequence for the synthesized objects. For example, the operation types of a triangular prism are represented as [5, 0, 1, 1, 1, 4, 6, 6, 6, 6].
[0118] Training Data Preparation. Two different approaches were employed for dataset preparation: 1) Random synthesis: CAD programs were randomly synthesized by assigning operation parameters within the predefined value range specified in Table 3. For example, when generating an add Une( • ) operation, a coordinate value of the endpoint from the range of [-1, 1] was randomly selected; 2) Synthesis with design rules: specific rules were established for parameter values. For instance, the extrusion depth of a circle was determined by the coordinates of its center point.
[0119] Training Details. The training dataset was divided into three sets, namely the training set, validation set, and test set, with a ratio of 8: 1 : 1, respectively. In Stage 1, the vector representations denoted as T were used to train the VAE for 1000 epochs using a latent dimension of 256. For Stage 2, the Enc( • ) was trained utilizing the pre-trained ResNetl8 model [6] for 50 epochs. The training was conducted using the Adam optimizer with a learning rate of 0.0001 and a batch size of 128.
[0120] Additional Examples
[0121] Table 4 introduces Fusion 360 CAD software python API as well as a domain specific language (DSL) developed based on it, namely Fusion 360 Gallery DSL [L], Table 5 gives a more comprehensive list of parameters used in all CAD operations. 7 parameters were used in the current study which are a subset of the parameters in Table 5.
[0122] Fusion 360 Gallery domain-specific language [L] (referred to herein as Gallery DSL) is used in the current study. Gallery DSL allows representing a design as a CAD program forming a simplified wrapper around the underlying Fusion 360 Python API. A CAD program consists of a sequence of Sketch and Extrude operation functions that iteratively modify the current geometry. Using the Gallery DSL, one can create novel objects by selecting a sketch plane and sketching on the plane using a sequence of curves (e.g., lines and arcs) to form closed-loop profiles. A profile can then be selected and extruded at a given distance and direction to form 3D objects.
[0123] Gallery DSL defines the current geometry (G) as a single global state that is updated with the sequence of operations. Gallery DSL can now only support two major types of CAD operations: Sketch and Extrude as highlighted in Table 4. (1) Sketch: A Sketch (S) operation is realized by the add sketch(I) operation, where l is a plane identifier on which the sketch will be created. I can be specified as a plane from the three canonical planes XY, YZ, and XZ or other planar faces (e.g., the side face of a cube) present in the current geometry G. After defining a sketch plane, a sequence of curves (Cu) including a Line (L), Arc (A), or Circle (C) can be drawn using add line(N, N, N, N), add arc(N, N, N, N, N), and add circle(N, N, N), respectively. L is specified by four numbers (N), representing the coordinates of the start and end points. A requires five numbers: the start point, the Arc’s center point, and the sweep angle. C is specified by three numbers with the first two representing the Circle’s center and the last one representing its radius. The sketch plane I in G is used as the reference coordinate system for specifying the coordinates. Executing a Sketch S operation results in enclosed regions (known as profiles) in the current geometry G. (1) Extrude: AnExtrude operation extrudes a profile from 2D into 3D which is realized by using add extrude(I, N, O), where I is the identifier for the profile, and N is a signed distance parameter defining the depth of extruding along the normal direction of the profile. The Boolean operation (O) specifies the behavior of the extruded 3D volume, e.g., add to or subtract from other 3D bodies.
[0124] The Gallery DSL cannot allow the creation of all kinds of objects for now but can still cover a vast range of them by using expressive sketches and extrude operations that have the built-in Boolean capability. It can also serve as a good starting point for supporting learning methods using the DSL. Nevertheless, it should be noted that the other CAD operations, such as revolve, sweep, and fillet, can be added to the Gallery DSL to expand its design grammar for a full-fledged CAD tool as shown in Table 4. The current study leveraged the current status of the Gallery DSL by only considering the Sketch and Extrude operations.Table 4: Gallery Domain Specific Language (current supported CAD operations are highlighted)Table 5: CAD parameters associated with Gallery DSL (parameters that are highlighted are used in the current study to formulate the vectorized design representation. For those whose value range is not specified indicates their associated CAD operations are not supported by Gallery DSL yet).
[0125] Trainins Data Preparation
[0126] Based on the data synthesis pipeline described in connection with Figure 3E, a collection of 5 template shapes (TS) were developed, as depicted in Figure 4A. Specifically, a sequence of operations 340b is established for each template shape 340a. These shapes are crafted using Sketch- Extrude operations, detailed in Table 1. To create the Sketch of each shape, Line, Arc, and Circle operations were utilized and the Extrude operation applied to generate the 3D volume of these shapes. TS 1-3 each correspond to unique sequences of operation types, while TS 4 and 5 are associated with three varied sequences. An example is provided for each template shape 340a.
[0127] Two distinct strategies were employed for dataset preparation: (1) Random generation of CAD programs (i.e., dataset without rules): random values are assigned to CAD operation parameters for the template sequences. This randomness is within predefined value ranges as outlined in Table 3. For example, for an add line ) operation, a number from the range [—1, 1] is randomly selected to determine the coordinates of endpoints. This method ensures diversity in the dataset by incorporating a wide range of parameter values, reflecting various potential CAD designs. (2) Generation of CAD programs with embedded design rules (i.e., dataset with rules): Contrary to the random approach, this method involves parameter selection based on predefined design rules. For example, the extrusion depth of a circle is determined by the coordinates of its center point. This strategy mimics the purpose-driven process typical of real-world design scenarios.
[0128] By incorporating design rules into data synthesis, design knowledge is embedded in the data. This method is expected to test how the embedded design principles would affect the effectiveness of CAD reconstruction from images. The combination ofrandom generation and rule-based design in dataset preparation allows for a comprehensive evaluation of the disclosed system’s capabilities in varying scenarios. Under each synthesis strategy, 2,000 shapes corresponding to every sequence type outlined in Figure 4A were synthesized, except for the sequence of TS 1, for which 6,000 shapes were synthesized. This was taken to ensure a balanced dataset in terms of both the length of sequences and the number of shapes for each template shape category. Consequently, this led to the creation of 22,000 CAD models for each strategy. These models were derived by processing the synthesized CAD programs through Fusion 360 software. For imaging purposes, all these 3D models were rendered using a uniform perspective camera positioned at (20.0,20.0,20.0) looking towards the origin (0.0, 0.0, 0.0) and all the images are in the resolution of 512x512 pixels. This process resulted in the image set X = {xk} =°00. The corresponding feature matrices Y = {yk}k=i°° were also saved during synthesis. The final training dataset was {xk>yk}k=i°0- Inaddition to a single group of CAD operations (e.g., “S, L, L, L, E” shown in FIG. 4A) this disclosure contemplates that the proposed method can be extended to complex geometries and assemblies and used to generate multiple groups of CAD operations. By way of example, “(0) S, L, L, L, E; (1) S, A, L, L, E; (2) S, L, A, L, E;..” can define a group of CAD operations, where (0), (1), (2), and so on represent the group number. In other words, in addition to outputting a sequence of operations 340b for template shapes (e.g., circle, arc), the method can be used to output CAD sequences for more complex shapes and geometries (e.g., a cylinder, a pyramid) defined by multiple groupings of CAD operations / sequences of operations 340b.
[0129] In Figure 4B, a schematic diagram illustrating the development of the Image2CADSeq network 400B is shown. The construction of the Image2CADSeq network 400B involved the application of target representation learning techniques, commonly employing target-embedding autoencoders (TEA) that utilize an autoencoder to obtain the latent representation of target objects
[0038] , In a recent development, Li et al.
[0010] introduced a target-embedding variational autoencoder (TEVAE) by extending the autoencoder to a variational autoencoder (VAE). This approach has demonstrated superior effectiveness in the cross-modal synthesis of 3D designs. However, their study did not include an empirical comparison between the two architectures in terms of the generated results. Therefore, as part of this disclosure’s contributions, two models were investigated — a baseline transformerbased AE and an improved version in the form of a transformer-based VAE for Stage 1.
[0130] Figure 4B shows the implementation of the Image2CADSeq network 400B. Two different encoder-decoder architectures were explored in Stage 1 302: (1) Baseline model: a transformer-based autoencoder (AE) 309, adapted from DeepCADNet [2], and (2) Enhanced model: a transformer-based variational autoencoder (VAE) 310, which extends the AE architecture. In Stage 2 304, the encoder is developed based on ResNetl8
[0046] , employing a dropout layer before the final layer to mitigate overfitting and enhance generalization.
[0131] The aim was to compare their performance and explore a better architecture for the Image2CADSeq network. The transformer-based AE 309, adapted from DeepCADNet [2], whose first and last layers were modified to capture and reconstruct the feature matrices 319a, 319b of the CAD programs in this study. Its primary objective is to learn and interpret the latent space of these matrices 319a, 319b, providing a robust foundation for accurate feature representation. The AE model 309 is used to establish a baseline for understanding and processing the complex structures inherent in CAD designs. It uses a typical reconstruction loss coupled with a regularization loss. Reconstruction loss ensures accurate reconstruction of input features in the output, while regularization loss prevents overfitting, promoting a more generalized model capable of handling various CAD designs.
[0132] In some implementations, as further depicted in Figure 4B, the transformerbased AE 309a includes a plurality of encoders (as shown, Encoder^, Encoder B. . .Encoder TV) configured in parallel such that each encoder is used to encode a different data group (e.g., an individual grouping of CAD operations that correspond with a portion of a complex shape or object). Then, the resulting embedding vector is fused before being fed into the decoder. Additionally, and / or alternatively, a single encoder can be modified to facilitate fusing of multiple embedding vectors, for example, using the first few layers of the encoder.
[0133] The transformer-based VAE'. Extending the AE architecture, the VAE 310 introduces a probabilistic approach to encoding, which is tailored to construct a smoother latent space, surpassing the AE 309 in terms of flexibility and adaptability. The VAE 310 employs a more complex loss function with KL-divergence loss, reconstruction loss, and regularization loss. The KL-divergence loss is pivotal in managing the probabilistic aspect of VAE 310, ensuring that the encoded distributions are effectively regularized. This, along with the reconstruction and regularization losses, forms a comprehensive approach to learning, capturing both the variance and the intricate details of CAD sequences. Stage 2 304 of the development incorporates an encoder based on ResNetl8 410
[0046] , In addition, a dropoutlayer 412 is positioned between the encoder 410 and the embedding vector layer 414 to prevent overfitting to the training data, thus maintaining its efficacy on unseen data. A regression loss was utilized between the embedding vector of the image and its corresponding latent vector obtained from the latent space in Stage 1 and a regularization loss to promote the generalizability of the Stage 2 encoder.
[0134] We divided each of the two training datasets, i.e., the datasets with and without rules, into three subsets: train, validation, and test set with a proportion of 8: 1 : 1. The validation set was used to monitor the training process, preventing the model from overfitting to the train set data and ensuring that the model’s generalization capabilities to the unseen data, e.g., test set data. A grid search strategy was employed in Stage 1 302 to find optimal hyperparameters of the neural network models and the training, aiming to minimize the training loss while maintaining good generalizability of the models. The AE 309 and VAE 310 models exhibited similar training trends, leading us to use the same hyperparameter set for both models. In conducted experiments, for Stage 1 302, a latent dimension size of 256 proved optimal for both models, resulting in the lowest reconstruction loss for the test set data among trials with dimensions of 64, 128, 256, and 512. Other hyperparameters include 500 epochs of training with a batch size of 512, the Adam optimizer, and a learning rate of 0.001. Moving to Stage 2, training was initiated with a pre-trained ResNetl8 model 410
[0046] that possesses a broad comprehension of various images. The S2 encoder was trained for 50 epochs using the Adam optimizer with a learning rate of 0.0001 and a batch size of 128. A dropout ratio of 0.4 was applied in the dropout layer 412.
[0135] Overall Evaluation of the CAD Programs
[0136] A study was conducted to compare two datasets in terms of sequence accuracy for CAD sequence generation. The dataset that incorporated design rules yielded significantly higher sequence accuracy compared to the dataset without rules (0.961 VS. 0.432) on the unseen test set data. This finding highlights the importance of integrating design rules into the data synthesis process, as it enhances data quality and subsequently improves the performance of the network. Additionally, the study analyzed the accuracy of operation parameters in relation to parameter tolerance, as depicted in Figures 5A-5D. With a tolerance of 3, the average parameter accuracy for all parameters, except t and I (for which tolerance was set to 0), was found to be 0.446. While it is impractical to have excessively large tolerance values, Figure 5A-5D demonstrates the Image2CADSeq network’s ability to predict the correct parameters varies for different parameters. Notably, the network exhibits poor performance in predicting the coordinates of the circle’s center. This issue is closelylinked to the design rules, as the study randomly selected a center point for circles during the data synthesis process.Table 6(b): Evaluation metrics for the CAD programsTables 6A and 6B: Comprehensive evaluation metrics for the image, the CAD program, and the 3D CAD model.
[0137] Figures 5E-5G provide a comparison of the Image2CADSeq_network’s performance in another study that was conducted, evaluated under two different architectures and two datasets. Specifically, there are three figures, each representing a unique combination of architecture and dataset. Figure 5E illustrates the performance of the network when employing the TEA architecture in conjunction with the no-rules dataset. In contrast, Figure 5F presents results derived from the same TEA architecture, but the network is trained on the with-rules dataset. A significant improvement was observed when employing the with-rules dataset for model training. Consequently, the TEVAE architecture was tested using the with- rules dataset only, and the results are presented in Figure 5G. Table 6 below provides comprehensive evaluation metrics for the image, the CAD programs, and the 3D CAD models. The figures illustrate the network’s performance across various metrics (as defined in Table 6) when applied to the first n operations in a CAD program.
[0138] N is limited to 6 to encompass the longest template sequences. According toFigure 4A, the maximum length of the template sequences is 5 for CAD operations inaddition to a non-CAD operation SOP, marking the start of a program. Typically, higher metric values indicate superior performance, except for the EDSOT, where lower values are preferable. To maintain a uniform direction of performance across all metrics and enhance the readability of the plotting, the EDSOT is presented in its negative form in the figures. Furthermore, when calculating metrics related to parameter accuracy, such as the accuracy of CAD programs (ACP) and the accuracy of parameterl (AP1), a tolerance level (q = 3) is introduced. This tolerance accounts for permissible deviations in the quantized continuous variables that have 256 levels but does not extend to discrete variables, such as the sketch plane identifier that has only 3 levels (refer to Table 3). The tolerance reflects the design problem’s criteria, allowing certain margins of error in parameter predictions that can be customized in different scenarios.
[0139] The metrics, classified into three hierarchical categories, Hl, H2, and H3 (see Table 6 (b)), include the accuracy of CAD programs (ACP), the accuracy of the sequence of operation types (ASOT), and the edit distance of the sequence of operation types (EDSOT) for Hl sequence evaluation; the accuracy of the operation types (AOT) and the accuracy of parameter (AP1) for the evaluation of the H2 sequence-based operation type; and the multiset similarity of operation types (MSOT) for the evaluation of the H3 set-based operation type. The results are analyzed according to this hierarchical structure. Upon analyzing the results of Figure 5E for the TEA trained using the no-rules dataset (referred to as Case 1), a downward trend was observed in all metrics as the sequence length increases. This is intuitive, and predicting longer sequences is inherently more difficult for the model. The ACP metric drops to zero at n = 3, indicating The ACP metric drops to zero at n = 3, indicating that the model struggles to accurately predict the entire CAD program including both the operation types and parameters even when the sequence is relatively short. Notably, the ACP’s decrease to roughly 0.4 at n = 2 suggests the model’s specific difficulty in predicting the sketch plane given the input images, because the second CAD operation — Sketch — defines the sketch plane’s position. In addition, both ASOT and the negative EDSOT metrics exhibit declines beginning with n = 3, together with the results of ACP, showing the limited sequence prediction capability of the model when the sequences get longer. Moreover, despite the model’s low values in Hl metrics and the low values in AP1of H2, it scores highly on the AOT metric of H2 and maintains high values of MSOT-TC and MSOT-CS in H3. This discrepancy and inconsistency probably result from the characteristics of the dataset, in which different shape categories share similar CAD sequences (see Figures 5E-5G).
[0140] For example, a ground truth (GT) sequence of TS 3, [“S”,“A”,“A”,“A”,“E”], could be mistakenly predicted as a TS 4 sequence, e.g., [“S”,“L”,“A”,“A”,“E”] by the model. Although such a prediction would be deemed as an incorrect sequence prediction when evaluated against ASOT, it would score well (i.e., 4 correct and 1 incorrect operation type) in terms of AOT due to the correctly predicted operation types. A dataset encompassing a broader array of design objects with diverse sequences of CAD operations might mitigate such discrepancies in the metric values, such as real-world designs collected from human designers. More significantly, it underscores the need for a comprehensive evaluation framework for image-to CAD sequence prediction, ensuring that models are thoroughly assessed from multiple perspectives. Otherwise, the performance of the models might be biased.
[0141] In Figure 5F, the performance of the TEA architecture trained on the dataset with rules (referred to as Case 2) is shown, in contrast to the earlier results in Figure 5E. In particular, this treatment improves significantly in most metrics, except for ACP and AP1. These exceptions, however, do not overshadow the overall enhancement in the model’s ability to predict sequences accurately. Specifically, while ACP does not show a significant improvement, its slower rate of decline at n = 2 and 3 indicates an improvement in the model’s performance of predicting sketch planes. Furthermore, although AP1does not show a significant improvement, it convergences to a higher value than the previous data treatment, suggesting an improved performance in parameter prediction.
[0142] The significant differences between Figures 5E and Figure 5F highlight the positive impact that design rules can have on the performance of the model inImage2CADSeq predictions. However, despite the improvement when including design rules, the model still faces challenges in accurately predicting parameters.
[0143] Building upon the insights from the first two cases, the model was evolved from TEA to TEVAE and trained using the dataset with rules (referred to as Case 3). The results of the TEVAE model are detailed in Figure 5G. It displays high accuracy in most metrics similar to the baseline performance of the TEA model from Case 2 but largely surpasses its performance in ACP and AP1. Particularly, the ACP metric shows a significant improvement in the TEVAE model and achieves a higher value at n = 3 (does not decreaseto zero as in Case 2). The AP1metric also reveals an upward trend, settling at a higher value than previously seen with the TEA model. For a more complete comparison of the three cases, the results of all metrics are summarized at n = 6 in Table 7 below.Table 7: Results for evaluation of first six operations of the CAD programs
[0144] This study demonstrates that the TEVAE model, when trained using the dataset with design rules, not only surpasses the TEA counterpart using the same dataset but also gives the best results across all metrics evaluated in all three cases.
[0145] Analysis on the overall parameter accuracy. The models demonstrate high accuracy in predicting the sequence of CAD operations but are less precise in parameter prediction. To facilitate a clearer comparison between the three cases with respect to parameter prediction accuracy, Figure 5H and Figure 51 are included to illustrate the relationship between parameter accuracy and tolerance using metrics ACP and AP1. Especially, in Figure 51 the blue dashed line with triangle markers (“Baseline w / o sketch paras”) represents the AP1value achieved by a hypothetical scenario of randomly guessing parameters given a specific tolerance, but without considering the Sketch parameter (i.e., the identifier of the sketch plane 7) whose values are not permitted for tolerance. The equation for the blue dashed line can be simplified to AP1= (— p2+ 511p + 256) / 65536. This line acts as a baseline to evaluate the model’s effectiveness in accurately predicting parameters if the sketch parameter is not considered.
[0146] To consider the Sketch parameter I, the characteristics of the dataset are taken into account (i.e., the ratio of each parameter taken among all possible parameters in the design representation of the CAD programs as outlined above). Accordingly, Equation (9) is plotted as the red dashed line with triangle markers in Figure 51. Note that the number of the sketch parameter takes 11 / 91 of all parameters. Additionally, a green dotted line is used to indicate the ideal scenario where the parameters are perfectly predicted with zero tolerance(i.e., GT). Other lines in Figure 5H and Figure 51 depict the corresponding metric values for different cases, providing a comprehensive view of the model’s performance in parameter prediction.
[0147] In both Figure 5H and Figure 51, the metric values increase with rising tolerance levels. A notable point in Figure 5H is that the ACP values for all three cases reach their highest at a tolerance of 255 and the corresponding values are 0.432, 0.961, and 0.967 for each case, in accordance with the ASOT values presented in Table 7. This can be interpreted as the result that when the entire CAD program is evaluated in terms of ACP given that all parameters are accurately predicted, ASOP is essentially being assessed. Additionally, a significant observation in Figure 5H and Figure 51 is how differently the models respond to changes in tolerance. Specifically, the TEVAE model, when trained using the dataset with rules, exhibits the highest sensitivity to changes in tolerance in contrast to cases 2 and 1. This trend suggests that the TEVAE model excels in parameter prediction compared to the TEA model. In Figure 51, a crucial observation is that all three lines exceed the baseline of the random guess of parameters. This indicates that the models are effectively learning parameter prediction from the training data. It is also important to note that the accuracy of these predictions depends on both the quality of the training data (for example, in this study, differentiated by the inclusion or exclusion of design rules) and the architecture of the model.
[0148] Analysis of Parameter Accuracy Based on Operation Types. To gain more insight into how the models perform in parameter prediction, the variation of AP1versus is plotted versus the tolerance for the operation parameters for each CAD operation type.
[0149] Figure 5J, Figure 5K, and Figure 5L are each a series of graphs depicting variations in AP1against tolerance levels (q = 0 -255) for specific parameters, corresponding to different CAD operations, Line, Circle, Arc, and Extrude, respectively. These results look into the model’s adaptability and accuracy across various CAD operations, providing a comprehensive understanding of its capabilities in different model architectures and datasets. Each graph includes a red dashed line representing the baseline as defined in Equation (8). In addition, the green dotted line illustrates the perfect prediction of the parameters with zero tolerance. The other lines show the AP1for specific parameters related to the respective CAD operations. To facilitate a more quantitative comparison of how well the parameters are predicted, the area under the curve (AUC) was computed for each parameter, as indicated in the upper right corner of each figure. Figure 5J, the result of Case 1, shows that the x, y co- coordinates of the center of the Circle, the sweep angle a of theArc, and the depth d of the Extrude align closely with the baseline. This suggests that the TEA model, when trained on a dataset without explicit design rules, performs similarly to random guessing for these specific parameters. However, the model still demonstrates the ability to learn certain patterns from the dataset, as evidenced by its recognition of the end- point of the Line, the radius of the Circle, and the endpoint of the Arc.
[0150] In Figure 5K, for Case 2, there is an evident improvement in all metric values compared to Case 1. This improvement highlights the enhanced ability of the TEA model to predict parameters. The significant distance of these values from the baseline indicates that the model has effectively learned the design rules embedded in the training data, enabling it to predict the corresponding parameters more effectively.
[0151] In Figure 5L, the results demonstrate an even better performance. All metric values not only surpass those in Figure 5K, but they also show a further deviation from the baseline, indicating a significant enhancement of the model’s predictive performance. These values are closely approaching the ground truth (GT) line, underscoring the refined ability of the TEVAE model to learn and apply the embedded design rules from the training data.
[0152] Overall Evaluation of the 3D CAD Models and Images. Figure 5M shows the summary of the parsing rate, intersection over union (loU), and mean squared error (MSE) outlined in Figure 4B of the three cases. An analysis of loU and MSE was performed by calculating the mean and standard deviation for each metric. Following this computation, the distribution of these values is depicted for each metric using a violin plot. These plots show the data distributions of the loU and MSE values move toward improved performance regions (i.e., higher values for loU and lower values for MSE). Aligning with observations in the evaluation of the CAD programs, as introduced in previous sections, the TEVAE trained using the dataset with design rules achieves the best performance among all three cases.
[0153] The Impact of Design Rules. The results in Figure 5M consistently show that the TEA model trained on the dataset with rules outperforms the one without rules. This result underscores the beneficial impact of design rules on the predictive performance of the model in Image2CADSeq tasks. This has three implications: (1) The dataset that incorporates design rules introduces a structured learning environment that guides the model to produce CAD designs that are not only more accurate but also more realistic. This suggests that the model benefits from learning embedded design principles or knowledge, not just raw data. (2) The inclusion of design rules can enhance the model’s generalizability. It also playsa critical role in minimizing the probability of creating unfeasible CAD designs. (3) Since piratical CAD designs often follow domain-specific standards and knowledge, data are therefore inevitably associated with rules. Thus, the proposed method has strong practical implications.
[0154] TEA vs. TEVAE. The improved predictive performance in TEVAE is due to the use of a variational autoencoder (VAE) that can capture the latent design representations of CAD programs. The effectiveness of the TEVAE architecture is demonstrated through improved performance across various metrics, significantly exceeding the results achieved in the TEA model. The superior performance of the TEVAE model can be largely attributed to the three advantages offered by VAEs. (1) Unlike traditional AEs, VAEs create a latent space that is well-structured and continuous. This design facilitates smoother interpolation between data points, enhancing the capture of meaningful variations in CAD designs. (2) The encoder in a VAE is more efficient in extracting relevant and prominent features from CAD programs than a standard AE. This efficiency stems from the VAE’s focus on capturing the underlying data distribution, rather than merely replicating input data. (3) The inclusion of the KL-divergence term in the VAE’s loss function helps reduce overfitting. It promotes the model to capture a broader data distribution rather than memorizing specific instances. This enhances TEVAE’ s generalizability on new, unseen data.
[0155] Artificial Intelligence and Machine Learning
[0156] The term “artificial intelligence” is defined herein to include any technique that enables one or more computing devices or computing systems (i.e., a machine) to mimic human intelligence. Artificial intelligence (Al) includes, but is not limited to, knowledge bases, machine learning, representation learning, and deep learning. The term “machine learning” is defined herein to be a subset of Al that enables a machine to acquire knowledge by extracting patterns from raw data. Machine learning techniques include, but are not limited to, logistic regression, support vector machines (SVMs), decision trees, Naive Bayes classifiers, and artificial neural networks. The term “representation learning” is defined herein to be a subset of machine learning that enables a machine to automatically discover representations needed for feature detection, prediction, or classification from raw data. Representation learning techniques include, but are not limited to, autoencoders and embeddings. The term “deep learning” is defined herein to be a subset of machine learning that enables a machine to automatically discover representations needed for feature detection, prediction, classification, etc. using layers of processing. Deep learning techniques include, but are not limited to, artificial neural network or multilayer perceptron (MLP).
[0157] Machine learning models include supervised, semi-supervised, and unsupervised learning models. In a supervised learning model, the model learns a function that maps an input (also known as feature or features) to an output (also known as target or targets) during training with a labeled data set (or dataset). In an unsupervised learning model, the model learns patterns (e.g., structure, distribution, etc.) within an unlabeled data set. In a semi-supervised model, the model learns a function that maps an input (also known as feature or features) to an output (also known as target or target) during training with both labeled and unlabeled data.
[0158] Artificial Neural Networks: An artificial neural network (ANN) is a computing system including a plurality of interconnected neurons (e.g., also referred to as “nodes”). This disclosure contemplates that the nodes can be implemented using a computing device (e.g., a processing unit and memory as described herein). The nodes can be arranged in a plurality of layers such as input layer, output layer, and optionally one or more hidden layers. An ANN having hidden layers can be referred to as deep neural network or multilayer perceptron (MLP). Each node is connected to one or more other nodes in the ANN. For example, each layer is made of a plurality of nodes, where each node is connected to all nodes in the previous layer. The nodes in a given layer are not interconnected with one another, i.e., the nodes in a given layer function independently of one another. As used herein, nodes in the input layer receive data from outside of the ANN, nodes in the hidden layer(s) modify the data between the input and output layers, and nodes in the output layer provide the results. Each node is configured to receive an input, implement an activation function (e.g., binary step, linear, sigmoid, tanH, or rectified linear unit (ReLU) function), and provide an output in accordance with the activation function. Additionally, each node is associated with a respective weight. ANNs are trained with a dataset to maximize or minimize an objective function. In some implementations, the objective function is a cost function, which is a measure of the ANN’S performance (e.g., error such as LI or L2 loss) during training, and the training algorithm tunes the node weights and / or bias to minimize the cost function. This disclosure contemplates that any algorithm that finds the maximum or minimum of the objective function can be used for training the ANN. Training algorithms for ANNs include, but are not limited to, backpropagation. It should be understood that an artificial neural network is provided only as an example machine learning model. This disclosure contemplates that the machine learning model can be any supervised learning model, semisupervised learning model, or unsupervised learning model. Optionally, the machine learningmodel is a deep learning model. Machine learning models are known in the art and are therefore not described in further detail herein.
[0159] A convolutional neural network (CNN) is a type of deep neural network that has been applied, for example, to image analysis applications. Unlike a traditional neural networks, each layer in a CNN has a plurality of nodes arranged in three dimensions (width, height, depth). CNNs can include different types of layers, e.g., convolutional, pooling, and fully-connected (also referred to herein as “dense”) layers. A convolutional layer includes a set of filters and performs the bulk of the computations. A pooling layer is optionally inserted between convolutional layers to reduce the computational power and / or control overfitting (e.g., by downsampling). A fully-connected layer includes neurons, where each neuron is connected to all of the neurons in the previous layer. The layers are stacked similar to traditional neural networks. GCNNs are CNNs that have been adapted to work on structured datasets such as graphs.
[0160] Logistic Regression: A logistic regression (LR) classifier is a supervised classification model that uses the logistic function to predict the probability of a target, which can be used for classification. LR classifiers are trained with a data set (also referred to herein as a “dataset”) to maximize or minimize an objective function, for example a measure of the LR classifier’s performance (e.g., error such as LI or L2 loss), during training. This disclosure contemplates that any algorithm that finds the maximum or minimum of the objective function can be used. LR classifiers are known in the art and are therefore not described in further detail herein.
[0161] Naive Bayes: A Naive Bayes’ (NB) classifier is a supervised classification model that is based on Bayes’ Theorem, which assumes independence among features (i.e., presence of one feature in a class is unrelated to presence of any other features). NB classifiers are trained with a data set by computing the conditional probability distribution of each feature given label and applying Bayes’ Theorem to compute conditional probability distribution of a label given an observation. NB classifiers are known in the art and are therefore not described in further detail herein.
[0162] KNN: A k-NN classifier is an supervised classification model that classifies new data points based on similarity measures (e.g., distance functions). k-NN classifier is a non-parametric algorithm, i.e., it does not make strong assumptions about the function mapping input to output and therefore has flexibility to find a function that best fits the data. k-NN classifiers are trained with a data set (also referred to herein as a “dataset”) by learning associations between all samples and classification labels in the training dataset. For example,k-NN classifiers can be trained to maximize or minimize a measure of the k-NN classifier’s performance during training. k-NN classifiers are known in the art and are therefore not described in further detail herein.
[0163] Ensemble: An majority voting ensemble is a meta-classifier that combines a plurality of machine learning classifiers for classification via majority voting. In other words, the majority voting ensemble’s final prediction (e.g., class label) is the one predicted most frequently by the member classification models. Majority voting ensembles are known in the art and are therefore not described in further detail herein.
[0164] Example Computing Device
[0165] Referring to Figure 6, an example computing device 600 upon which the methods described herein may be implemented is illustrated. It should be understood that the example computing device 600 is only one example of a suitable computing environment upon which the methods described herein may be implemented. Optionally, the computing device 600 can be a well-known computing system including, but not limited to, personal computers, servers, handheld or laptop devices, multiprocessor systems, microprocessorbased systems, network personal computers (PCs), minicomputers, mainframe computers, embedded systems, and / or distributed computing environments including a plurality of any of the above systems or devices. Distributed computing environments enable remote computing devices, which are connected to a communication network or other data transmission medium, to perform various tasks. In the distributed computing environment, the program modules, applications, and other data may be stored on local and / or remote computer storage media.
[0166] In its most basic configuration, computing device 600 typically includes at least one processing unit 606 and system memory 604. Depending on the exact configuration and type of computing device, system memory 604 may be volatile (such as random access memory (RAM)), non-volatile (such as read-only memory (ROM), flash memory, etc.), or some combination of the two. This most basic configuration is illustrated in Figure 6 by dashed line 602. The processing unit 606 may be a standard programmable processor that performs arithmetic and logic operations necessary for the operation of the computing device 600. The computing device 600 may also include a bus or other communication mechanism for communicating information among various components of the computing device 600.
[0167] Computing device 600 may have additional features / functionality. For example, computing device 600 may include additional storage such as removable storage 608 and non-removable storage 610, including, but not limited to, magnetic or optical disksor tapes. Computing device 600 may also contain network connect! on(s) 616 that allow the device to communicate with other devices. Computing device 600 may also have input device(s) 614 such as a keyboard, mouse, touch screen, etc. Output device(s) 612, such as a display, speakers, printer, etc., may also be included. The additional devices may be connected to the bus in order to facilitate the communication of data among the components of the computing device 600. All these devices are well-known in the art and need not be discussed at length here.
[0168] The processing unit 606 may be configured to execute program code encoded in tangible, computer-readable media. Tangible, computer-readable media refers to any media that is capable of providing data that causes the computing device 600 (i.e., a machine) to operate in a particular fashion. Various computer-readable media may be utilized to provide instructions to the processing unit 606 for execution. Example of tangible, computer- readable media may include, but is not limited to, volatile media, non-volatile media, removable media, and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules or other data. System memory 604, removable storage 608, and non-removable storage 610 are all examples of tangible, computer storage media. Examples of tangible, computer-readable recording media include, but are not limited to, an integrated circuit (e.g., field-programmable gate array or application-specific IC), a hard disk, an optical disk, a magneto-optical disk, a floppy disk, a magnetic tape, a holographic storage medium, a solid- state device, RAM, ROM, electrically erasable program read-only memory (EEPROM), flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices.
[0169] In an example implementation, the processing unit 606 may execute program code stored in the system memory 604. For example, the bus may carry data to the system memory 604, from which the processing unit 606 receives and executes instructions. The data received by the system memory 604 may optionally be stored on the removable storage 608 or the non-removable storage 610 before or after execution by the processing unit 606.
[0170] It should be understood that the various techniques described herein may be implemented in connection with hardware or software or, where appropriate, with a combination thereof. Thus, the methods and apparatuses of the presently disclosed subject matter, or certain aspects or portions thereof, may take the form of program code (i.e., instructions) embodied in tangible media, such as floppy diskettes, CD-ROMs, hard drives, orany other machine-readable storage medium where, when the program code is loaded into and executed by a machine, such as a computing device, the machine becomes an apparatus for practicing the presently disclosed subject matter. In the case of program code execution on programmable computers, the computing device generally includes a processor, a storage medium readable by the processor (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device. One or more programs may implement or utilize the processes described in connection with the presently disclosed subject matter, e.g., through the use of an application programming interface (API), reusable controls, or the like. Such programs may be implemented in a high-level procedural or object-oriented programming language to communicate with a computer system. However, the program(s) can be implemented in assembly or machine language, if desired. In any case, the language may be a compiled or interpreted language, and it may be combined with hardware implementations.
[0171] Discussion
[0172] Computer-aided design (CAD) tools empower designers to craft and modify 3D models through a series of CAD operations, commonly referred to as a CAD sequence. In scenarios where digital CAD files are not accessible, reverse engineering (RE) has been used for reconstructing 3D CAD models. Recent advances have seen the rise of data-driven approaches for RE, with a primary focus on converting 3D data, such as point clouds, into 3D models in boundary representation (B-rep) format. However, obtaining 3D data poses significant challenges, and B-rep models do not reveal knowledge and insights into the process of modeling the 3D models. To that end, the research described herein introduces a novel data-driven method, the Image2CADSeq model. This innovative approach is designed to reverse engineer CAD models by processing images as input and generating CAD sequences. These sequences can then be translated into B-rep models using a solid modeling kernel. Distinct from B-rep models, CAD sequences offer enhanced flexibility for modifying individual steps of the model creation, providing a deeper understanding of the historical construction process of the CAD model. To rigorously evaluate the Image2CADSeq model, a comprehensive evaluation framework was developed. The model was trained on a specially synthesized dataset, and various network architectures were explored to optimize performance. The results have been highly promising, demonstrating the model’s proficiency in generating CAD sequences from a single image input. However, refining the accuracy of parameters remains an area for future enhancement and research.
[0173] Computer-aided design (CAD) systems can significantly decrease design time by avoiding the need for labor-intensive manual drawings traditionally required [1], Contemporary CAD systems such as Fusion 360, and SOLIDWORKS enable designers to create and modify CAD models through a sequence of CAD operations. CAD models are structured, parametric, or operation-based 2D or 3D design, allowing for a high level of control and flexibility in the design process. They are particularly suited for industrial and engineering designs where precision and the ability to easily modify designs are crucial. There are two major types of CAD models: 1) constructed solid geometry (CSG) and 2) parametric CAD models including CAD sequence data and boundary representation (B-rep) models. In contrast, discrete 3D representations, such as meshes and point clouds are more static and less flexible in terms of parametric editing and design exploration [2, 3], However, in certain scenarios, the CAD model of a product may not be readily available due to various factors, including outdated documentation and the lack of digital records. Reverse engineering (RE) is employed to overcome these obstacles, utilizing measurement and analysis tools to reconstruct CAD models [2], Integrating RE with CAD systems allows designers to leverage the advantages of existing products while incorporating their own innovative ideas and improvements. However, manual reconstruction of CAD models is labor-intensive and time-consuming, leading to the exploration of data-driven approaches, such as converting 3D point clouds into CAD models [3], [4], Compared to point clouds or voxels, images are simpler to acquire given, for example, the popularity of mobile devices. In addition, in certain cases, there is no access to 3D models to get these 3D data and only 2D images may be available. Thus, the question arises: How can images be reverse engineered into CAD sequences to allow CAD designers to easily interpret and edit the CAD models during the modeling process? After a thorough literature review, it was noted that there is a scarcity of research on reconstructing 3D CAD models from images. The objective is to develop a computational framework that can generate a sequence of CAD operations based on a single image (referred to as “Image2CADSeq”). A CAD sequence offers advantages over 3D CAD models, enabling flexibility in geometry modification, and facilitating a better understanding of the historical process of the CAD model construction. The novelties and contributions of the proposed approach are summarized as follows, a). To the best of the inventors’ knowledge, this study is the first attempt at predicting a CAD sequence of operation sequences given a single image input (i.e., image-to-CAD sequence prediction). A target-embedding variational autoencoder (TEVAE) architecture
[0010] is developed for this design problem. A high level of accuracy and efficiency in generating CAD models from 2Dimages was achieved. The proposed method has the potential to greatly streamline the CAD design process, reducing the time and effort required to create 3D models from 2D sketches or photographs, b). A novel data synthesis pipeline was created based on the design grammars defined in the Fusion 360 Gallery domain-specific language (DSL). The pipeline can generate synthetic data that closely resembles real-world images and CAD models. It can be used as a method for data augmentation to improve the quality and quantity of existing training data, providing a more diverse and robust dataset for training image2CADSeq models, c). A robust evaluation framework was designed to thoroughly assess the Image2CADSeq task, encompassing three key components: the CAD sequence, the CAD model, and the corresponding images. Specifically for the CAD sequence evaluation, an innovative three-hierarchy, two-level evaluation system was developed. This system is meticulously crafted to provide a comprehensive analysis of the CAD sequence prediction, ensuring a multi-faceted examination from diverse perspectives.
[0174] The proposed method has the potential to revolutionize existing CAD systems by making the CAD model reconstruction process more accessible. This would enable both experienced and novice designers to actively contribute to the design, promoting design collaboration and design education. Moreover, it has the potential to provide a unique pathway to involve end users in the design process, fostering design democratization.
[0175] The main objective of the Image2CADSeq network is to generate a CAD sequence, consisting of operation types and associated parameters, based on an input image. For training purposes, the study utilized synthesized data representing simple shape primitives. Although the parameter prediction accuracy is less than 50%, the network demonstrates a promising accuracy of up to 96% in predicting the sequence of operation types. This indicates the network’s potential to accurately predict the CAD sequence given a single image.
[0176] The exemplary system and method has the potential to bring about transformative changes in existing CAD systems, revolutionizing the product development cycle. For example, a designer can merely generate a first order approximation of a design as a sketch or image to which a CAD, CAM, or CAE model with parametric lists can be generated for subsequent modification. The generated vectorized map or matrix provides a comprehensive list of CAD-software specific operations that is compatible and native of existing CAD software.
[0177] Moreover, it has the capacity to democratize the CAD model reconstruction process, allowing individuals with limited experience or expertise to actively participate. Byremoving barriers, it can also facilitate customer engagement in design activities, promoting the democratization of design.
[0178] Conclusion
[0179] Various sizes and dimensions provided herein are merely examples. Other dimensions may be employed.
[0180] Although example embodiments of the present disclosure are explained in some instances in detail herein, it is to be understood that other embodiments are contemplated. Accordingly, it is not intended that the present disclosure be limited in its scope to the details of construction and arrangement of components set forth in the following description or illustrated in the drawings. The present disclosure is capable of other embodiments and of being practiced or carried out in various ways.
[0181] It must also be noted that, as used in the specification and the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise. Ranges may be expressed herein as from “about” or “ 5 approximately” one particular value and / or to “about” or “approximately” another particular value. When such a range is expressed, other exemplary embodiments include from the one particular value and / or to the other particular value.
[0182] By “comprising” or “containing” or “including” is meant that at least the name compound, element, particle, or method step is present in the composition or article or method, but does not exclude the presence of other compounds, materials, particles, method steps, even if the other such compounds, material, particles, method steps have the same function as what is named.
[0183] In describing example embodiments, terminology will be resorted to for the sake of clarity. It is intended that each term contemplates its broadest meaning as understood by those skilled in the art and includes all technical equivalents that operate in a similar manner to accomplish a similar purpose. It is also to be understood that the mention of one or more steps of a method does not preclude the presence of additional method steps or intervening method steps between those steps expressly identified. Steps of a method may be performed in a different order than those described herein without departing from the scope of the present disclosure. Similarly, it is also to be understood that the mention of one or more components in a device or system does not preclude the presence of additional components or intervening components between those components expressly identified.
[0184] The term “about,” as used herein, means approximately, in the region of, roughly, or around. When the term “about” is used in conjunction with a numerical range, itmodifies that range by extending the boundaries above and below the numerical values set forth. In general, the term “about” is used herein to modify a numerical value above and below the stated value by a variance of 10%. In one aspect, the term “about” means plus or minus 10% of the numerical value of the number with which it is being used. Therefore, about 50% means in the range of 45%-55%. Numerical ranges recited herein by endpoints include all numbers and fractions subsumed within that range (e.g., 1 to 5 includes 1, 1.5, 2, 2.75, 3, 3.90, 4, 4.24, and 5).
[0185] Similarly, numerical ranges recited herein by endpoints include subranges subsumed within that range (e.g., 1 to 5 includes 1-1.5, 1.5-2, 2-2.75, 2.75-3, 3-3.90, 3.90-4, 4-4.24, 4.24-5, 2-5, 3-5, 1-4, and 2-4). It is also to be understood that all numbers and fractions thereof are presumed to be modified by the term “about.”
[0186] The following patents, applications, and publications, as listed below and throughout this document, describes various application and systems that could be used in combination the exemplary system and are hereby incorporated by reference in their entirety herein.[1] Rosato, D., and Rosato, D., 2003. “5 - computer-aided design”. In Plastics Engineered Product Design, D. Rosato and D. Rosato, eds. Elsevier Science, Amsterdam, pp. 344-380.[2] Wu, R., Xiao, C., and Zheng, C., 2021. “Deepcad: A deep generative network for computer-aided design models”. In Proceedings of the IEEE / CVF International Conference on Computer Vision, pp. 6772-6782.[3] Para, W., Bhat, S., Guerrero, P., Kelly, T., Mitra, N., Guibas, L. J., and Wonka, P., 2021. “Sketchgen: Generating constrained cad sketches”. Advances in Neural Information Processing Systems, 34.[4] Varady, T., Martin, R. R., and Cox, J., 1997. “Reverse engineering of geometric models — an introduction”. Computer-aided design, 29(4), pp. 255-268.[5] Buonamici, F., Carfagni, M., Furferi, R., Govemi, L., Lapini, A., and Volpe, Y., 2018. “Reverse engineering modeling methods and tools: a survey”. Computer-Aided Design and Applications, 15(3), pp. 443-464.[6] Uy, M. A., Chang, Y.-Y., Sung, M., Goel, P., Lambourne, J. G., Birdal, T., and Guibas, L. J., 2022.“Point2cyl: Reverse engineering 3d objects from point clouds to extrusion cylinders”. InProceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 11850-11860.[7] Ren, D., Zheng, J., Cai, J., Li, J., and Zhang, J., 2022. “Extrudenet: Unsupervised inverse sketch-and-extrude for shape parsing”. In European Conference on Computer Vision, Springer, pp. 482-498.[8] Lamboume, J. G., Willis, K., Jayaraman, P. K., Zhang, L., Sanghi, A., and Malekshan, K. R., 2022. “Reconstructing editable prismatic cad from rounded voxel models”. In SIGGRAPH Asia 2022 Conference Papers, pp. 1-9.[9] Li, P., Guo, J., Zhang, X., and Yan, D.M., 2023. “Secadnet: Self-supervised cad reconstruction by learning sketch-extrude operations”. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 16816-16826.
[0010] Li, X., Xie, C., and Sha, Z., 2022. “A predictive and generative design approach for three-dimensional mesh shapes using target-embedding variational autoencoder”. Journal of Mechanical Design, 144(11), p. 114501.
[0011] Shi, Z., Peng, S., Xu, Y., Geiger, A., Liao, Y., and Shen, Y., 2022. “Deep generative models on 3d representations: A survey”. arXiv preprint arXiv:2210.15663.
[0012] Li, X., Wang, Y ., and Sha, Z., 2023. “Deep learning methods of cross-modal tasks for conceptual design of product shapes: A review”. Journal of Mechanical De- sign, 145(4), p. 041401.
[0013] Ren, D., Zheng, J., Cai, J., Li, J., Jiang, H., Cai, Z., Zhang, J., Pan, L., Zhang, M., Zhao, H., and Yi, S., 2021. “Csg-stump: A learning friendly csg-like representation for interpretable shape parsing”. 2021 IEEE / CVF International Conference on Computer Vision (ICCV), null, pp. 12458-12467.
[0014] Sharma, G., Goyal, R., Liu, D., Kalogerakis, E., and Maji, S., 2017. “Csgnet: Neural shape parser for constructive solid geometry”. 2018 IEEE / CVF Conference on Computer Vision and Pattern Recognition, null, pp. 5515-5523.
[0015] Sharma, G., Goyal, R., Liu, D., Kalogerakis, E., and Maji, S., 2019. “Neural shape parsers for constructive solid geometry”. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44, pp. 2628-2640.
[0016] Kania, K., Zieba, M., and Kajdanowicz, T., 2020. “Ucsg-net-unsupervised discovering of constructive solid geometry tree”. Advances in Neural Information Processing Systems, 33, pp. 8776-8786.
[0017] Willis, K. D., Jayaraman, P. K., Lamboume, J. G., Chu, H., and Pu, Y., 2021. “Engineering sketch generation for computer-aided design”. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 2105-2114.
[0018] Seff, A., Zhou, W., Richardson, N., and Adams, R. P., 2021. “Vitruvion: A generative model of parametric cad sketches”. In International Conference on Learning Representations.
[0019] Ganin, Y., Bartunov, S., Li, Y., Keller, E., and Saliceti, S., 2021. “Computer-aided design as language”. Advances in Neural Information Processing Systems, 34.
[0020] Yang, Y., and Pan, H., 2022. “Discovering design concepts for cad sketches”. ArXiv, abs / 2210.14451, p. null.
[0021] Wang, X., Xu, Y., Xu, K., Tagliasacchi, A., Zhou, B., Mahdavi-Amiri, A., and Zhang, H., 2020. “Pie- net: Parametric inference of point cloud edges”. Advances in neural information processing systems, 33, pp. 20167-20178.
[0022] Sharma, G., Liu, D., Maji, S., Kalogerakis, E., Chaudhuri, S., and Meeh, R., 2020. “Parsenet: A parametric surface fitting network for 3d point clouds”. In Computer Vision- ECCV 2020: 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part VII 16, Springer, pp. 261-276.
[0023] Smirnov, D., Bessmeltsev, M., and Solomon, J., 2020. “Learning manifold patchbased representations of man-made shapes”. In International Conference on Learning Representations.
[0024] Guo, H., Liu, S., Pan, H., Liu, Y., Tong, X., and Guo, B., 2022. “Complexgen: Cad reconstruction by b- rep chain complex generation”. ACM Transactions on Graphics (TOG), 41(4), pp. 1-18.
[0025] Jayaraman, P. K., Lamboume, J. G., Desai, N., Willis, K., Sanghi, A., and Morris, N. J., 2022. “Solidgen: An autoregressive model for direct b-rep synthesis”. Transactions on Machine Learning Research.
[0026] Wang, K., Zheng, J., and Zhou, Z., 2022. “Neural face identification in a 2d wireframe projection of a manifold object”. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 1622-1631.
[0027] Xu, X., Willis, K. D., Lamboume, J. G., Cheng, C - Y ., Jayaraman, P. K., and Furukawa, Y., 2022. “Skexgen: Autoregressive generation of cad construction sequences with disentangled codebooks”. In International Conference on Machine Learning, PMLR, pp. 24698- 24724.
[0028] Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I., 2017. “Attention is all you need”. Advances in neural information processing systems, 30.
[0029] Willis, K. D., Pu, Y., Luo, J., Chu, H., Du, T., Lamboume, J. G., Solar-Lezama, A., and Matusik, W., 2021. “Fusion 360 gallery: A dataset and environment for programmaticcad construction from human design sequences”. ACM Transactions on Graphics (TOG), 40(4), pp. 1-24.
[0030] Xu, X., Peng, W ., Cheng, C.-Y., Willis, K. D., and Ritchie, D., 2021. “Inferring cad modeling sequences using zone graphs”. In Proceedings of the IEEE / CVF conference on computer vision and pattern recognition, pp. 6062-6070.
[0031] Li, C., Pan, H., Bousseau, A., and Mitra, N. J., 2020. “Sketch2cad: Sequential cad modeling by sketching in context”. ACM Transactions on Graphics (TOG), 39(6), pp. 1- 14.
[0032] Li, C., Pan, H., Bousseau, A., and Mitra, N. J., 2022. “Free2cad: Parsing freehand drawings into cad com- mands”. ACM Transactions on Graphics (TOG), 41(4), pp. 1-16.
[0033] Hinton, G. E., and Salakhutdinov, R. R., 2006. “Reducing the dimensionality of data with neural networks”, science, 313(5786), pp. 504-507.
[0034] Kingma, D. P., and Welling, M., 2014. “Auto-encoding variational bayes.”. In International Conference on Learning Representations, Y. Bengio and Y. LeCun, eds.
[0035] Bengio, Y., Courville, A., and Vincent, P., 2013. “Representation learning: A review and new perspectives”. IEEE transactions on pattern analysis and machine intelligence, 35(8), pp. 1798-1828.
[0036] Khastavaneh, H., and Ebrahimpour-Komleh, H., 2019. “Representation learning techniques: An overview”. In The 7th International Conference on Contemporary Issues in Data Science, Springer, pp. 89-104.
[0037] Gomari, D. P., Schweickart, A., Cerchietti, L., Paietta, E., Fernandez, H., Al-Amin, H., Suhre, K., and Krumsiek, J., 2022. “Variational autoencoders learn transferrable representations of metabolomics data”. Communications Biology, 5(1), p. 645.
[0038] Jarrett, D., and van der Schaar, M., 2020. “Target embedding autoencoders for supervised representation learning”. In International Conference on Learning Representations.
[0039] Le, L., Patterson, A., and White, M., 2018. “Supervised autoencoders: Improving generalization performance with unsupervised regularizes”. Advances in neural information processing systems, 31.
[0040] Girdhar, R., Fouhey, D. F., Rodriguez, M., and Gupta, A., 2016. “Learning a predictable and generative vector representation for objects”. In European Conference on Computer Vision, Springer, pp. 484-499.
[0041] Mostajabi, M., Maire, M., and Shakhnarovich, G., 2018. “Regularizing deep networks by modeling and predicting label structure”. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5629-5638.
[0042] Dalca, A. V., Guttag, J., and Sabuncu, M. R., 2018. “Anatomical priors in convolutional networks for unsupervised biomedical segmentation”. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 9290-9299.
[0043] Carlier, A., Danelljan, M., Alahi, A., and Timofte, R., 2020. “Deepsvg: A hierarchical generative network for vector graphics animation”. Advances in Neural Information Processing Systems, 33, pp. 16351-16361.
[0044] Navarro, G., 2001. “A guided tour to approximate string matching”. ACM computing surveys (CSUR), 33(1), pp. 31-88.
[0045] Bajusz, D., R'acz, A., and H'eberger, K., 2015. “Why is tanimoto index an appropriate choice for fingerprint based similarity calculations?”. Journal of cheminformatics, 7(1), pp. 1-13.
[0046] He, K., Zhang, X., Ren, S., and Sun, J., 2016. “Deep residual learning for image recognition”. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770-778.
[0047] Song, B., Zurita, N. S., Zhang, G., Stump, G., Balon, C., Miller, S., Yukish, M., Cagan, J., and McComb, C., 2020. “Toward hybrid teams: A platform to understand humancomputer collaboration during the design of complex engineered systems”. In Proceedings of the Design Society: DESIGN Conference, Vol. 1, Cambridge University Press, pp. 1551— 1560.[1*] Rosato, D., and Rosato, D., 2003. “5 - computer-aided design”. In Plastics Engineered Product Design, D. Rosato and D. Rosato, eds. Elsevier Science, Amsterdam, pp. 344-380. [2*] Buonamici, F., Carfagni, M., Furferi, R., Governi, L., Lapini, A., and Volpe, Y., 2018. “Reverse engineering modeling methods and tools: a survey”. Computer-Aided Design and Applications, 15(3), pp. 443-464.[3*] Wu, R., Xiao, C., and Zheng, C., 2021. “Deepcad: A deep generative network for computer-aided design models”. In Proceedings of the IEEE / CVF International Conference on Computer Vision, pp. 6772-6782.[4*] Uy, M. A., Chang, Y.-Y., Sung, M., Goel, P., Lambourne, J. G., Birdal, T., and Guibas, L. J., 2022. “Point2cyl: Reverse engineering 3d objects from point clouds to extrusioncylinders”. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 11850-11860.[5*] Li, X., Xie, C., and Sha, Z., 2022. “A predictive and generative design approach for three-dimensional mesh shapes using target-embedding variational autoencoder”. Journal of Mechanical Design, 144(11), p. 114501.[6*] He, K., Zhang, X., Ren, S., and Sun, J., 2016. “Deep residual learning for image recognition”. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770-778.[7*] Willis, K. D., Pu, Y., Luo, J., Chu, H., Du, T., Lambourne, J. G., Solar-Lezama, A., and Matusik, W., 2021. “Fusion 360 gallery: A dataset and environment for programmatic cad construction from human design sequences”. ACM Transactions on Graphics (TOG), 40(4), pp. 1-24.[8*] Carlier, A., Danelljan, M., Alahi, A., and Timofte, R., 2020. “Deepsvg: A hierarchical generative network for vector graphics animation”. Advances in Neural Information Processing Systems, 33, pp. 16351-16361.[L] Willis, K. D. D., Pu, Y., Luo, J., Chu, H., Du, T., Lambourne, J. G., Solar-Lezama, A., and Matusik, W., 2021, “Fusion 360 Gallery: A Dataset and Environment for Programmatic Cad Construction from Human Design Sequences,” ACM Trans. Graph., 40(4), pp. 1-24.
Claims
What is claimed is:
1. A method for generating a parametric list for a computer-aided design (CAD) application comprising: providing at least one image of an object; generating, using a trained machine learning model and based on the at least one image, a vectorized object corresponding with the object; generating, using the trained machine learning model, the parametric list based on the vectorized object, wherein the parametric list includes a sequence of predefined operations and associated parameters to construct a corresponding CAD object, and wherein the parametric list is used, and is modifiable, in a CAD software or CAD workflow to construct a three-dimensional (3D) model of the CAD object.
2. The method of claim 1, wherein the trained machine learning model is trained using a training data set of images and corresponding vectorized objects and / or sequences of CAD operations.
3. The method of any one of claims 1-2, wherein the machine learning model is trained in a first training stage and a second training stage.
4. The method of claim 3, wherein the first training stage comprises unsupervised learning, and wherein the second training stage comprises supervised learning.
5. The method of any one of claims 1-4, wherein the trained machine learning model comprises a target-embedding variational autoencoder (TEVAE) architecture.
6. The method of claim 5, wherein the TEVAE architecture comprises a first encoderdecoder sub-model or variational autoencoder (VAE) with transformer-based blocks that is configured to encode sequence information of the predefined operations and their corresponding parameters into a latent space.
7. The method of claim 6, wherein the TEVAE architecture comprises a second encoder sub-model that is configured to regress the latent space using images (i.e., input feature objects).
8. The method of any one of claims 1-7, wherein the parametric list comprises between 6 and 30 discrete and / or continuous parameters.
9. The method of claim 8, wherein each parameter corresponds with a user-specified predefined value range.
10. The method of any one of claims 8-9, wherein the parameters comprise one or more of a operation type, an identifier of a sketch plane, at least one end coordinate point x, at least one end coordinate point y, a sweep or resolve angle, a radius, an identifier of profile, an identifier of reference, an extrusion distance, Boolean operations, a scale factor, a Mode, a Rho value and a width.
11. The method of any one of claims 1-10, wherein the CAD object is employed in computer-aided engineering (CAE) analysis or computer-aided manufacturing (CAM) analysis.
12. The method of any one of claims 1-11, wherein the CAD object is employed in a generative artificial intelligence model.
13. A system comprising: a processor; and a memory having instructions thereon, wherein the instructions when executed by the processor, cause the processor to: receive at least one image of an object; generate, using a trained machine learning model and based on the at least one image, a vectorized object corresponding with the object; generate, using the trained machine learning model, a parametric list based on the vectorized object, wherein the parametric list includes a sequence of predefined operations to construct a corresponding CAD object, and wherein the parametric list is used, and is modifiable, in a CAD software or CAD workflow to construct a 3D model of the CAD object.
14. The system of claim 13, wherein the instructions when executed by the processor, cause the processor to: train the machine learning model using a training data set of images and corresponding vectorized objects and / or sequences of CAD operations.
15. The system of claim 13 or 14, wherein the instructions when executed by the processor, cause the processor to: train the machine learning model in a first training stage and a second training stage.
16. The system of claim 15, wherein the first training stage comprises unsupervised learning, and wherein the second training stage comprises supervised learning.
17. The system of any one of claims 13-16, wherein the trained machine learning model comprises a target-embedding variational autoencoder (TEVAE) architecture.
18. The system of claim 17, wherein the TEVAE architecture comprises a first encoderdecoder sub-model or variational autoencoder (VAE) with transformer-based blocks that is configured to encode sequence information of the predefined operations and their corresponding parameters into a latent space.
19. The system of claim 18, wherein the TEVAE architecture comprises a second encoder sub-model that is configured to regress the latent space using images.
20. A non-transitory computer readable medium comprising a memory having instructions stored thereon, which when executed by a processor, cause the processor to: receive at least one image of an object; generate, using a trained machine learning model and based on the at least one image, a vectorized object corresponding with the object; generate, using the trained machine learning model, a parametric list based on the vectorized object, wherein the parametric list includes a sequence of predefined operations to construct a corresponding CAD object, and wherein the parametric list is used, and is modifiable, in a CAD software or CAD workflow to construct a 3D model of the CAD object.
Citation Information
Patent Citations
Method and system for generating a geometric component using machine learning models
US20230252207A1
Machine learning-based generation of constraints for computer-aided design (CAD) assemblies
US20230267248A1
Method and system for providing a three-dimensional computer aided-design (CAD) model in a CAD environment
WO2022039741A1
Cited By
Multi-source data fusion CAD drawing method and system based on deep learning
CN120876645A
Tea technology large model construction method and system, computer equipment and medium
CN120893546A
A method, system, computer equipment, and media for constructing a large-scale tea technology model.
CN120893546B
Brain-like continuous learning method and device inspired by degeneracy structure
CN121257614A