Deep Learning Generative Models
A deep learning generative model trained on 3D objects with feature-penalizing loss functions addresses the challenge of generating functionally appropriate 3D modeled objects, enhancing design and manufacturing efficiency by ensuring geometric and functional accuracy.
Patent Information
- Application Number
- JP2021156903
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-09-25
- Filing Date
- 2021-09-27
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2041-09-27
AI Technical Summary
Existing solutions for generating 3D modeled objects, particularly representing machine parts or assemblies, lack the ability to accurately output objects that meet functional appropriateness and structural integrity, focusing predominantly on geometric or structural features without considering interdependencies and physical interactions.
A deep learning generative model is trained using a dataset of 3D modeled objects, incorporating feature scores that penalize deviations from geometric, structural, and functional descriptors, such as stability, durability, and interaction affordances, utilizing variational autoencoders or generative adversarial networks to minimize loss and ensure accurate output.
The model generates 3D modeled objects that are functionally valid and accurate, reflecting intended use characteristics by minimizing feature score deviations, improving ergonomics and accuracy in design and manufacturing processes.
Smart Images

Figure 0007782998000020 
Figure 0007782998000021 
Figure 0007782998000022
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to the field of computer programs and systems, and more particularly to methods, devices, and programs related to deep learning generative models that output 3D modeled objects, each representing a mechanical part or an assembly of mechanical parts. [Background technology]
[0002] Many systems and programs are available on the market for designing, engineering, and manufacturing objects. CAD is an acronym for computer-aided design, which refers to software solutions for designing objects. CAE is an acronym for computer-aided engineering, which refers to software solutions for simulating the physical behavior of future products. CAM is an acronym for computer-aided manufacturing, which refers to software solutions for defining manufacturing processes and operations. In such computer-aided design systems, the graphical user interface plays a key role in the efficiency of the technology. These technologies can be incorporated within product lifecycle management (PLM) systems. PLM refers to a business strategy in which companies share product data, apply common processes, and leverage enterprise knowledge to support the development of products from conception to the end of their lifecycle, across the extended enterprise concept. PLM solutions offered by Dassault Systèmes (under the trademarks CATIA, ENOVIA, and DELMIA) provide an engineering hub that organizes product engineering knowledge, a manufacturing hub that manages manufacturing engineering knowledge, and an enterprise hub that enables enterprise integration and connectivity to both the engineering and manufacturing hubs. Overall, the system provides an open object model that links products, processes, and resources, enabling dynamic, knowledge-based product creation and decision support to drive optimized product definition, manufacturing preparation, production, and service. In this and other contexts, deep learning and especially deep learning generative models for 3D modeling objects has become widely important.
[0003] The following documents are relevant to this field and are referenced below: [1]Oliver Van Kaick, Hao Zhang, Ghassan Hamarneh, and Daniel Cohen-Or. A survey on shape correspondence. CGF, 2011. [2]Johan WH Tangelder and Remco C Veltkamp. A survey of content based 3d shape retrieval methods. Multimedia tools and applications, 39(3): 441-471, 2008. [3]Wayne E Carlson. An algorithm and data structure for 3d object synthesis using surface patch intersections. In SIGGRAPH, 1982. [4] J. Wu, C. Zhang, T. Xue, WT Freeman, JB Tenenbaum. Learning a Probabilistic Latent Space of Object Shapes via 3D Generative-Adversarial Modeling, 2016. [5]Siddhartha Chaudhuri, Daniel Ritchie, Jiajun Wu, Kai Xu, Hao Zhang. Learning Generative Models of 3D Structures. EUROGRAPHICS 2020. [6]R. B. Gabrielsson, V. Ganapathi-Subramanian, P. Skraba, and L. J. Guibas. Topology-Layer for Machine Learning, 2019. [7]Erwin Coumans and Yunfei Bai. Pybullet. A python module for physics simulation for games, robotics and machine learning. http: / / pybullet.org, 2016-2019. [8]Rundi Wu, Yixin Zhuang, Kai Xu Hao Zhang, Baoquan Chen. PQ-NET: A Generative Part Seq2Seq Network for 3D Shapes. CVPR 2020. [9]Qi, C. R., Su, H., Mo, K., & Guibas, L. J. PointNet: Deep learning on point sets for 3D classification and segmentation. CVPR, 2017.
[0004] [1, 2, 3] concern methods for obtaining and synthesizing new shapes by borrowing and combining parts from an existing patrimony of 3D objects (e.g., datasets). Generative methods demonstrate the possibility of synthesizing 3D modeled objects (i.e., coming up with new content rather than simply combining and obtaining existing heritage), thereby proposing conceptually novel objects. Traditional generative methods focus exclusively on either the geometric or structural features of the 3D modeled object and learn and apply generative models of the 3D content [4, 5]. [4] introduces a model that is a generative decoder that maps noise samples (e.g., uniform or Gaussian distribution) onto a voxel grid as an output with a distribution close to that of the data generation distribution. [5] investigates learning 3D generative models that recognize the structure of the data, e.g., the arrangement and relationships between parts of the 3D modeled object's shape.
[0005] However, there is still a need for improved solutions when it comes to outputting 3D modeled objects, each representing a machine part or an assembly of machine parts. Summary of the Invention
[0006] Accordingly, a computer-implemented method for training a deep learning generative model that outputs 3D modeled objects, each representing a machine part or an assembly of machine parts, is provided. The training method includes providing a dataset of 3D modeled objects. Each 3D modeled object represents a machine part or an assembly of machine parts. The training method further includes training the deep learning generative model based on the dataset. Training the deep learning generative model includes minimizing a loss. The loss includes, for each output 3D modeled object, a term that penalizes one or more feature scores of the respective 3D modeled object. Each feature score measures the degree of non-respect of a respective feature descriptor among one or more feature descriptors by the machine part or assembly of machine parts.
[0007] The training method may include one or more of the following steps: The loss further includes, for each output 3D modeled object, another term that penalizes the discrepancy of the shape of the respective 3D modeled object relative to the dataset. The other terms are the reconstruction loss between each 3D modeled object and the corresponding ground truth 3D modeled object in the dataset, the adversarial loss associated with the dataset, or Mapping distance, which measures the shape dissimilarity between each 3D modeled object and the corresponding modeled object in the dataset Includes: Deep learning generative models include 3D generative neural networks. 3D generative neural networks include variational autoencoders or generative adversarial networks. Deep learning generative models are a 3D generative neural network, the 3D generative neural network including a variational autoencoder, the other terms including a reconstruction loss and a variational loss; a 3D generative neural network, including a generative adversarial network, where the other terms include an adversarial loss; or A mapping model followed by a 3D generative neural network, where the 3D neural network is pre-trained and other terms include a mapping distance, and the 3D generative neural network optionally includes a variational autoencoder or a generative adversarial network. It consists of one of the following: The training includes, for each 3D modeled object, calculating a feature score for the 3D modeled object, the calculation including: deterministic functions, Simulation-based engines, or Deep Learning Functions This is done by applying one or more of: The calculation is performed by a deep learning function, the deep learning function being trained on another dataset, the other dataset including 3D objects each associated with a respective feature score, the respective feature score being deterministic functions, Simulation-based engines, or Deep Learning Functions The calculation is performed by using one or more of: The one or more feature descriptors include a connectivity descriptor. One or more feature descriptors are one or more geometric descriptors, and / or One or more affordances Includes: The one or more geometric descriptors are: a physical stability descriptor, for a machine part or assembly of machine parts, that describes the stability of the machine part or assembly of machine parts under the application of gravity alone; and / or A durability descriptor for a mechanical part or assembly of mechanical parts, which describes the ability of the mechanical part or assembly of mechanical parts to withstand the application of gravity and external mechanical forces. Includes: One or more affordances can be a support affordance descriptor, for a mechanical part or assembly of mechanical parts, that describes the ability of the mechanical part or assembly of mechanical parts to withstand only the application of an external mechanical force; and / or a drag coefficient descriptor, for a mechanical part or assembly of mechanical parts, that describes the effect of a fluid environment on the mechanical part or assembly of mechanical parts; a containment affordance descriptor, for a mechanical part or an assembly of mechanical parts, that describes the ability of the mechanical part or the assembly of mechanical parts to contain other objects in an interior volume of the mechanical part; a hold affordance descriptor, for a mechanical part or assembly of mechanical parts, that describes the ability of the mechanical part or assembly of mechanical parts to support other objects via a mechanical connection; and / or a hanging affordance descriptor, for a mechanical part or assembly of mechanical parts, that describes the ability of the mechanical part or assembly of mechanical parts to be supported via a mechanical connection; Includes: Each 3D modeled object in the dataset is A piece of furniture, Electric vehicles, Non-motorized vehicles, or tool Represents.
[0008] Additionally, a method of using a deep learning generative model that outputs 3D modeled objects each representing a machine part or assembly of machine parts and that has been trained (i.e., trained) according to the training method is provided. The method of use includes providing a deep learning generative model and applying the deep learning generative model to output one or more 3D modeled objects each representing a respective machine part or assembly of machine parts.
[0009] Additionally, a machine learning process is provided, including a method for training and then using the machine learning process.
[0010] Additionally, computer programs including instructions for carrying out training methods, methods of use, and / or machine learning processes are provided.
[0011] Further provided is a device including a data storage medium having a program and / or a trained deep learning generative model recorded thereon. The device may form a non-transitory computer-readable medium. Alternatively, the device may include a processor coupled to the data storage medium. Thus, the device may form a system. The system may further include a graphical user interface coupled to the processor. [Brief explanation of the drawings]
[0012] Embodiments of the present disclosure will now be described, by way of non-limiting examples, with reference to the accompanying drawings, in which: [Figure 1] FIG. 1 is a flowchart of a training method. [Figure 2] FIG. 2 is a diagram showing an example of a graphical user interface of the system. [Figure 3] FIG. 3 is a diagram illustrating an example of a system. [Figure 4] FIG. 4 illustrates a training method, method of use, and / or machine learning process. [Figure 5] FIG. 5 illustrates a training method, method of use, and / or machine learning process. [Figure 6] FIG. 6 illustrates a training method, method of use, and / or machine learning process. [Figure 7] FIG. 7 illustrates a training method, method of use, and / or machine learning process. [Figure 8] FIG. 8 illustrates a training method, method of use, and / or machine learning process. [Figure 9] FIG. 9 illustrates a training method, method of use, and / or machine learning process. [Figure 10] FIG. 10 illustrates a training method, method of use, and / or machine learning process. [Figure 11] FIG. 11 illustrates a training method, method of use, and / or machine learning process. [Figure 12] FIG. 12 illustrates a training method, method of use, and / or machine learning process. [Figure 13] FIG. 13 illustrates a training method, method of use, and / or machine learning process. [Figure 14] FIG. 13 illustrates a training method, method of use, and / or machine learning process. [Figure 15] FIG. 13 illustrates a training method, method of use, and / or machine learning process. [Figure 16] FIG. 13 illustrates a training method, method of use, and / or machine learning process. DETAILED DESCRIPTION OF THE INVENTION
[0013] Referring to FIG. 1, a computer-implemented method for training a deep learning generative model is provided. The deep learning generative model is configured to output 3D modeled objects. Each 3D modeled object represents a mechanical part or an assembly of mechanical parts. The training method includes providing a dataset of 3D modeled objects (S10). Each 3D modeled object represents a mechanical part or an assembly of mechanical parts. The training method further includes training the deep learning generative model based on the dataset (S20). The training (S20) includes minimizing a loss. The loss includes, for each output 3D modeled object, a term that penalizes one or more feature scores of the respective 3D modeled object. Each feature score measures the degree of non-respect of a respective feature descriptor among one or more feature descriptors by the mechanical part or assembly of mechanical parts.
[0014] The training method constitutes an improved solution for outputting 3D modeled objects each representing a machine part or an assembly of machine parts.
[0015] In particular, the training method can obtain a generative model configured to output a 3D modeled object. The generative model can be used to automatically generate (i.e., output or synthesize) one or more 3D modeled objects, thereby improving ergonomics in the field of 3D modeling. Furthermore, the training method obtains a generative model via training S20 based on the dataset provided in S10, where the generative model is a deep learning generative model. Thus, the training method leverages improvements provided by the field of machine learning. In particular, the deep learning generative model can learn to replicate the data distribution of diverse 3D modeled objects present in the dataset and output the 3D modeled object based on the learning.
[0016] Furthermore, the deep learning generative model outputs 3D modeled objects that are particularly accurate with respect to the intended use or objective characteristics of the represented mechanical part or assembly of mechanical parts within the intended context. This is thanks to the training (S20) process, which includes minimizing a specific loss. The loss specifically includes, for each output 3D modeled object, a term that penalizes one or more feature scores for each 3D modeled object. When this specific term is included in the loss, the deep learning generative model is trained to output 3D objects that meet the functional appropriateness measured by one or more feature scores. In fact, each feature score measures the degree to which the mechanical part or assembly of mechanical parts respects a respective feature descriptor among one or more feature descriptors. In an example, such feature descriptors may include any cues related to the shape, structure, or any type of interaction between the represented mechanical part or assembly of mechanical parts and other objects or mechanical forces within the intended context. Thus, the training (S20) process may evaluate the functional appropriateness of the 3D modeled objects of the dataset within the intended context, for example, with respect to the stability, physical feasibility, and quality of interaction of the shape or structure within the intended context. Thus, deep learning generative models output functionally valid and particularly accurate 3D modeled objects.
[0017] Additionally, training S20 may investigate functional features of the 3D modeled objects of the dataset. In an example, in addition to the geometric and structural features of the 3D objects, training S20 may investigate interdependencies among the geometric, structural, topological, and physical features of the 3D objects. Furthermore, training S20 may be used to output 3D modeled objects that refine or modify the functionality of each 3D modeled object of the dataset, thereby improving the accuracy of the output 3D modeled objects for the dataset.
[0018] Any method herein is computer-implemented. This means that all steps of the training method (including S10 and S20) and all steps of the use method are performed by at least one computer or any system. Thus, the method steps are performed by a computer, possibly fully automatically or semi-automatically. In an example, triggering of at least some steps of the method may be performed through user-computer interaction. The level of user-computer interaction required depends on the expected level of automation and may balance the need to implement user preferences. In an example, this level may be user-defined and / or predefined.
[0019] A typical example of a computer implementation of the methods herein is performing the methods using a system adapted for this purpose. The system may include a processor coupled to a memory and a graphical user interface (GUI), the memory having recorded thereon a computer program including instructions for performing the method. The memory may also store a database. The memory is any hardware adapted for such storage, possibly including several physically distinct parts (e.g., one for the program and optionally one for the database).
[0020] A modeled object is any object defined by data stored, for example, in a database. Thus, the term "modeled object" refers to the data itself. Depending on the type of system, the modeled object may be defined by various types of data. A system may, in fact, be any combination of a CAD system, a CAM system, a PDM system, and / or a PLM system. In these different systems, the modeled object is defined by the corresponding data. Thus, we may speak of CAD objects, PLM objects, PDM objects, CAE objects, CAM objects, CAD data, PLM data, PDM data, CAM data, and CAE data. However, these systems are not exclusive of one another, as a modeled object may be defined by data corresponding to any combination of these systems. Thus, a system may be both a CAD system and a PLM system.
[0021] A CAD system refers to any system at least adapted for designing a modeled object based on a graphical representation of the modeled object, such as CATIA. In this case, data defining the modeled object includes data enabling the representation of the modeled object. The CAD system may provide a CAD representation of the modeled object, for example, using edges or lines, and possibly faces or surfaces. The modeled object specifications may be stored in a single CAD file or multiple CAD files. Typical sizes of files representing modeled objects in a CAD system are in the range of one megabyte per part. A modeled object may typically be an assembly of thousands of parts.
[0022] In the context of CAD, a modeled object may typically be a 3D modeled object. "3D modeled object" means any object modeled by data that allows for its 3D representation. The 3D representation allows the part to be viewed from any angle. For example, when a 3D modeled object is represented in 3D, it can be manipulated and rotated around any of its axes or around any axis of the screen on which the representation is displayed.
[0023] Any 3D modeled object herein, including any 3D modeled object output by a deep learning generative model, may represent the shape of a (e.g., mechanical) part or assembly of parts (i.e., an assembly of parts may be viewed as the part itself from the perspective of the training method, and thus the assembly of parts, or training method may be applied to each part of the assembly individually), or more generally, the shape of a product manufactured in the real world, such as any rigid body assembly (e.g., a mobile mechanism). Thus, any of the above 3D modeled objects may represent an industrial product which may be any machine part such as a part of a motorized or non-motorized ground vehicle (including, for example, automobiles and light truck equipment, racing cars, motorcycles, trucks and motor equipment, trucks and buses, trains), a part of an aircraft (including, for example, aircraft equipment, aerospace equipment, propulsion equipment, defense products, aviation equipment, space equipment), a part of a naval vehicle (including, for example, naval equipment, commercial vessels, offshore equipment, yachts and workboats, marine equipment), a general machine part (including, for example, industrial manufacturing machinery, large mobile machinery or equipment, installation equipment, industrial equipment products, metal fabrication products, tire manufacturing products), an electric machine or electronic component (including, for example, consumer electronics products, security and / or control and / or instrumentation products, computing and communications equipment, semiconductors, medical devices and instruments), a consumer product (including, for example, furniture, home and garden products, leisure products, fashion products, durable goods retailer products, non-durable goods retailer products), packaging (including, for example, food and beverage and tobacco, beauty and personal care, household goods packaging), etc. Any of the above 3D modeled objects, including any 3D modeled objects output by a deep learning generative model, may then be integrated with, for example, a CAD software solution or system as part of a virtual design that enables the subsequent design of products in a variety of unlimited industry sectors, including aerospace, architecture, construction, consumer goods, high-tech devices, industrial equipment, transportation, marine, and / or offshore oil and gas production or transportation.
[0024] A PLM system further refers to any system adapted to the management of modeled objects that represent physically manufactured products (or products to be manufactured). In a PLM system, the modeled objects are therefore defined by data that are suitable for the manufacture of the physical objects. These may typically be dimensional and / or tolerance values. For the correct manufacture of the objects, it is indeed better to have such values.
[0025] CAM solution further refers to a solution that is software on hardware adapted to manage the manufacturing data of a product. Manufacturing data typically includes data related to the product to be manufactured, the manufacturing process, and the resources required. CAM solutions are used to plan and optimize the entire manufacturing process of a product. For example, they can provide CAM users with information on the feasibility, duration of the manufacturing process, or the number of resources, such as specific robots, that can be used in a particular step of the manufacturing process, thus enabling decisions regarding management or necessary investments. CAM is a subsequent process after the CAD process and potentially the CAE process. Such CAM solutions are offered by Dassault Systèmes under the DELMIA® trademark.
[0026] FIG. 2 shows an example GUI of the system, which can be used to view and / or design (e.g., edit) any 3D modeled object output by a deep learning generative model and integrated as part of a virtual design.
[0027] The GUI 2100 may be a typical CAD-like interface, with standard menu bars 2110, 2120 and bottom and side toolbars 2140, 2150. Such menus and toolbars include a set of user-selectable icons, each associated with one or more operations or functions, as known in the art. Some of these icons are associated with software tools adapted for editing and / or working with the 3D modeled object 2000 displayed in the GUI 2100. The software tools may be grouped into workbenches. Each workbench includes a subset of software tools. In particular, one of the workbenches is an editing workbench, suitable for editing geometric features of the modeled product 2000. During operation, the designer may, for example, pre-select a portion of the object 2000 and then initiate an operation (e.g., changing dimensions, color, etc.) or edit geometric constraints by selecting the appropriate icon. For example, a typical CAD operation is modeling punching or folding of a 3D modeled object displayed on the screen. The GUI may, for example, display data 2500 related to the displayed product 2000. In the illustrated example, the data 2500 displayed as a "feature tree" and its 3D representation 2000 relate to a brake assembly including a brake caliper and disc. The GUI may further exhibit various types of graphic tools 2130, 2070, 2080, for example, to facilitate 3D orientation of objects, trigger simulations of manipulation of the edited product, or render various attributes of the displayed product 2000. A cursor 2060 may be controlled by a haptic device to allow the user to interact with the graphic tools.
[0028] FIG. 3 shows an example of a system, where the system is a client computer system, e.g., a user's workstation.
[0029] The client computer in this example includes a central processing unit (CPU) 1010 connected to an internal communication BUS 1000 and a random access memory (RAM) 1070 also connected to the BUS. The client computer further includes a graphical processing unit (GPU) 1110 associated with a video random access memory 1100 connected to the BUS. The video RAM 1100 is also known in the art as a frame buffer. A mass storage controller 1020 manages access to mass memory devices such as a hard drive 1030. Mass memory devices suitable for embodying computer program instructions and data include all forms of non-volatile memory, including, by way of example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, magnetic disks such as internal hard disks and removable disks, magneto-optical disks, and CD-ROM disks 1040. Any of the foregoing may be supplemented by, or incorporated into, specially designed ASICs (application-specific integrated circuits). A network adapter 1050 manages access to a network 1060. The client computer may also include a cursor control device, a keyboard, or other haptic device 1090. A cursor control device is used in the client computer to allow a user to selectively position a cursor at any desired position on the display 1080. Furthermore, the cursor control device allows the user to select various commands and input control signals. The cursor control device includes several signal generating devices for inputting control signals to the system. Typically, the cursor control device may be a mouse, and the buttons on the mouse are used to generate the signals. Alternatively or additionally, the client computer system may include a sensitive pad and / or a sensitive screen.
[0030] A computer program may include computer-executable instructions, including means for causing the system to perform any of the methods described above. The program may be recordable on any data storage medium, including the system's memory. The program may be implemented, for example, in digital electronic circuitry, or computer hardware, firmware, software, or a combination thereof. The program may be implemented as an apparatus, for example, an article embodied in a machine-readable storage device for execution by a programmable processor. The method steps may be performed by a programmable processor executing a program of instructions that performs the functions of the method by manipulating input data and generating output. Thus, the processor may be programmable and coupled to receive and transmit data and instructions from and to a data storage system, at least one input device, and at least one output device. The application program may be implemented in a high-level procedural or object-oriented programming language, or in assembly or machine language as appropriate. In either case, the language may be compiled or interpreted. The program may be a full installation program or an update program. The application of the program on a system in any case provides instructions for performing any of the methods described herein.
[0031] A deep learning generative model trained according to the training method may be part of a process for designing a 3D modeled object. "Designing a 3D modeled object" refers to an action or series of actions that is at least part of a process for creating a 3D modeled object. For example, the process may include applying a deep learning generative model to output one or more 3D modeled objects and, optionally, implementing one or more design changes to each output 3D modeled object. The process may include displaying a graphical representation of each output 3D modeled object, and the optional one or more design changes may include applying CAD operations using CAD software, for example, via user graphical interaction.
[0032] Designing may involve using a deep learning generative model trained according to the training method to create a 3D modeled object from scratch, thus improving the ergonomics of the design process. Indeed, a user may not need to use cumbersome tools to create a 3D modeled object from scratch and may instead focus on other tasks in the design process. Alternatively, designing may involve providing a (previously created) 3D modeled object to a deep learning generative model trained according to the training method. The deep learning generative model may map the input modeled object to an output 3D modeled object, and subsequent design is performed on the output 3D modeled object of the deep learning generative model. The output 3D modeled object is more accurate in terms of functionality than the input 3D modeled object. In other words, the deep learning generative model may be used to modify or improve / optimize the input 3D modeled object, resulting in an improved design process, functionally speaking. The input 3D modeled object may be geometrically realistic but functionally invalid or suboptimal. In this case, the usage method may improve this.
[0033] The design process may be included in a manufacturing process, which may include, after the design process is performed, manufacturing a physical product corresponding to the modeled object. In either case, the modeled object designed by this method may represent a manufacturing object. Thus, the modeled object may be a modeled solid (i.e., a modeled object representing a solid). The manufacturing object may be a product, such as a part, or an assembly of parts. Because deep learning generative models improve the design process of creating modeled objects, they also improve the manufacturing of the product, thus increasing the productivity of the manufacturing process.
[0034] A "deep learning model" refers to any data structure representing a set of computations, at least a portion of which may be trained based on a dataset of 3D modeled objects. The training may be performed by a set of techniques known in the field of machine learning, particularly deep learning. A trained deep learning model, i.e., a deep learning model trained based on a dataset of 3D modeled objects, is "generative," i.e., configured to generate and output one or more 3D objects. In examples, the deep learning generative model may include a variational autoencoder or a generative adversarial network, and / or any other neural network (e.g., a convolutional neural network, a recurrent neural network, a discriminative model). In such examples, the deep learning generative model may generate a family of synthesized 3D objects with at least some accuracy in terms of their geometric shape.
[0035] The loss minimization may be performed in any manner, such as gradient descent or any other minimization known in the art. For each output 3D modeled object, the loss includes a term that penalizes one or more feature scores of the respective 3D modeled object.
[0036] "Feature score" means any metric (e.g., a number, a vector) attributable to each object that measures the degree of non-respect of each feature descriptor among one or more feature descriptors by a mechanical part or assembly of mechanical parts. By convention, a lower feature score may indicate a better match (or respect) to each feature descriptor.
[0037] A "functional descriptor" refers to any functional attribute of a 3D modeled object, i.e., a function that characterizes (i.e., represents) the physical function of the represented real-world object (mechanical part or assembly of mechanical parts). Physical function refers to the quality (of a real-world object) of being functional, i.e., the ability of the represented real-world object to perform or be capable of performing within a physical context of use, i.e., its use or purpose. A physical function can be a physical property that the represented real-world object possesses and that can be affected by physical interaction. Such physical properties can include, for example, the represented real-world object's electrical conductivity, thermal conductivity, drag in a fluid environment, density, stiffness, elastic deformation, tensile strength, and resilience. A physical interaction can be an interaction with another mechanical object or objects, or the presence of an environmental stimulus that affects the physical properties of the represented real-world object. An environmental stimulus can be any physical input (to the represented object) coming from at least a portion of the intended environment, e.g., a heat source from the environment or the environment including fluid flow. A physical function can be a mechanical function. Mechanical functionality may be properties of spatial structure and mechanical stability that a represented real-world object possesses and that may be affected under mechanical interaction. Properties of spatial structure and mechanical stability of a represented object may include the connectivity of the represented real-world object's components, the stability of the represented real-world object in use, its durability, or any affordances of the represented real-world object when affected by mechanical interaction. The mechanical interaction may include any mechanical interaction between the represented object and other mechanical objects and / or external mechanical stimuli. Examples of mechanical interactions between the represented object and other mechanical objects may include collision (e.g., including any type of impact with other objects and / or any type of friction with other objects), conflict with other objects within an enclosed volume, supporting other objects, holding other objects, or being suspended via a mechanical connection.The mechanical interaction of the represented object may include mechanical constraints of motion on the represented object, e.g., static or dynamic motion constraints. The mechanical interaction may further include any mechanical response of the represented object to an external mechanical stimulus. The external mechanical stimulus may consist of the application of a mechanical force to the represented object, e.g., a force, stress, or vibration force that changes from a resting state to a moving state. A functional descriptor may include multiple physical or mechanical functions, which may be combined as a one-dimensional descriptor or a multidimensional vector.
[0038] Thus, the feature score quantifies the degree to which the represented real-world object has physical functionality relevant to its use or purpose, as expressed by each of the one or more feature descriptors. The loss thus penalizes output 3D modeled objects that do not sufficiently match the characteristics of the represented real-world object's use or purpose. The loss thus enables learning to explore interdependencies between one or more physical and / or mechanical features of the represented real-world object characterized by one or more feature descriptors, for example, between the placement, structure, and physical and / or mechanical interactions of the object within its intended context. As a result, a generative model trained according to the training method outputs a 3D modeled object that is a particularly accurate representation of the corresponding represented mechanical part or assembly of mechanical parts, and that, in terms of its mechanical functionality, is, for example, geometrically stable, physically realizable, and has improved interaction qualities with the intended context of use.
[0039] The loss may further include another term for each output 3D modeled object that penalizes the inconsistency of the shape of the respective 3D modeled object with respect to the dataset. "Inconsistency of the shape of each 3D modeled object with respect to the dataset" refers to the inconsistency or mismatch between the shape of the respective 3D modeled object and the shape of at least some elements of the dataset. The inconsistency or mismatch may consist of the difference or distance (exact or approximate) between the shape of the respective 3D modeled object and the shape of at least some elements of the dataset. Thus, the learning explores the interdependency between the features (captured by penalizing the feature scores) and the shapes of the 3D modeled objects, thereby outputting accurate and realistic 3D modeled objects whose shapes best match the features of the represented mechanical part or assembly of mechanical parts.
[0040] Other terms may include a reconstruction loss, an adversarial loss, or a mapping distance. Thus, training may focus on improving shape consistency through either a reconstruction loss, an adversarial loss, or a mapping distance. Other terms, including the reconstruction loss between each 3D modeled object and its corresponding ground truth 3D modeled object in the dataset, may consist of a term that penalizes geometric dissimilarity between each 3D modeled object and its corresponding ground truth 3D modeled object in the dataset. Thus, a generative model trained according to the provided method may be used for improved 3D shape reconstruction that respects the intended function. The adversarial loss associated with the dataset may include or consist of a term that minimizes the discrepancy between the distribution of the dataset and the distribution of the generated 3D output modeled object. Thus, minimizing the discrepancy improves the shape consistency of the generated 3D output object with respect to the distribution of the dataset. The mapping distance measures the shape dissimilarity between each 3D modeled object and its corresponding modeled object in the dataset. Thus, the generative model outputs a 3D modeled object that may be more accurate (at least in terms of its functionality and shape consistency) with respect to the corresponding modeled object in the dataset.
[0041] The deep learning generative model may include a 3D generative network. A "3D generative neural network" means a deep learning generative model that forms a neural network, where the neural network is trainable according to a machine learning-based optimization, and all trainable elements of the neural network (e.g., all weights of the neural network) are trained together (e.g., during a single machine learning-based optimization). Thus, the deep learning generative model may utilize such a 3D generative neural network to improve aspects of the accuracy of the output 3D modeled object.
[0042] 3D generative neural networks can include variational autoencoders or generative adversarial networks (e.g., classical generative adversarial networks, or latent generative adversarial networks followed by a 3D converter). Such 3D generative neural networks improve the accuracy of the output 3D model.
[0043] A variational autoencoder consists of two parts: an encoder and a decoder. The encoder receives an input and outputs a distribution probability. The decoder reconstructs the input given samples from the distribution output by the encoder. For example, the distribution can be set to Gaussian, so that the encoder outputs two vectors of the same size representing, for example, the mean and variance of the distribution probability, and the decoder reconstructs the input based on the distribution. A variational autoencoder is trained by jointly training the encoder and decoder via a variational loss and a reconstruction loss.
[0044] A generative adversarial network (GAN) consists of two networks: a generator and a discriminator. The generator receives low-dimensional latent variables sampled from a Gaussian distribution as input. The generator's output can be the same type of data in a training dataset (e.g., a classic GAN) or the same type of data in a latent space (e.g., a latent GAN). The generator can be trained using a discriminator trained to perform binary classification on its input between two classes: "real" or "fake." The input must be classified as "real" if it is from the training dataset and as "fake" if it is from the generator. During the training phase, while the discriminator is trained to perform its binary classification task, the generator is trained to "fool" the discriminator by generating samples that are classified as "real" by the discriminator. To jointly train both networks, the GAN can be trained through an adversarial loss.
[0045] Alternatively, the 3D generative neural network may be from a family of 3D generative neural networks, including hybrid generative networks that can be built based on variational autoencoders and generative adversarial networks.
[0046] The deep learning generative model may consist of a 3D generative neural network, and the learning method thus trains all elements of the 3D generative neural network during machine learning-based optimization.
[0047] In such cases, the 3D generative neural network may include a variational autoencoder, and the other terms may include a reconstruction loss and a variational loss. Thus, a deep learning generative model trained according to the training method jointly trains the encoder and decoder of the variational autoencoder while minimizing a feature score. Thus, the deep learning generative model outputs a functionally valid 3D modeled object while leveraging the advantages of the variational autoencoder. Such advantages include, for example, sampling and interpolating two latent vector representations from a latent space (or performing other types of arithmetic on the latent space), and outputting a functional 3D modeled object based on the interpolation by a decoder trained according to the training method.
[0048] Instead of a variational autoencoder, the 3D generative neural network can include a generative adversarial network, and other terms can include adversarial losses. The adversarial losses can be of any kind, such as discriminator loss, minimax loss, non-saturation loss, etc. Therefore, the deep learning generative model can output functionally valid 3D modeled objects while leveraging the advantages of a generative adversarial network.
[0049] A deep learning generative model trained according to this method can be used within a method for synthesizing 3D modeled objects (i.e., creating a 3D modeled object from scratch), which can then be integrated as part of a design process, particularly within a CAD process, thereby providing users of the CAD system with accurate and realistic 3D modeled objects while improving the ergonomics of the design process. The synthesis can include obtaining latent representations from the latent space of the trained 3D generative neural network and generating the 3D modeled object from the obtained latent representations. The latent representations for generating the 3D modeled object can be obtained from any type of sampling from the latent space, for example, as a result of performing arithmetic operations or as a result of interpolation between at least two sampled latent representations from the latent space. In a further example, the latent representations can be obtained as a result of providing an input 3D modeled object to a deep learning generative model. Thus, the deep learning generative model can output a 3D modeled object with improved functionality over the input 3D modeled object.
[0050] Instead of a deep learning generative model composed of a 3D generative neural network, the deep learning generative model may be composed of a mapping model followed by a 3D generative neural network. In such a case, the 3D generative neural network may be pre-trained, and the other term may include a mapping distance. In such an alternative, the 3D generative neural network may optionally include a variational autoencoder or a generative adversarial network. The mapping model may be trained by a mapping distance that penalizes shape dissimilarity between each output 3D modeled object and a corresponding modeled object in the dataset. The corresponding modeled object in the dataset may have been sampled or may correspond to (and thus added to) the provided 3D modeled object. Thus, the provided 3D modeled object may have been output from another generative model or as a result of a previous design process. Thus, the training method may explore the interdependencies between functionality and the shape of the 3D modeled object to output a 3D modeled object with improved functionality. Thus, a deep learning generative model may focus on providing a particularly accurate output 3D modeled object with improved functionality and shape consistency while leveraging the generative advantages of a 3D generative network, such as synthesizing a 3D modeled object from the latent space of the 3D generative neural network, or may output a 3D modeled object with improved functionality over a provided 3D modeled object, which may already have leveraged improvements because, for example, it was part of a previous design process or has been output by another generative model.
[0051] A deep learning generative model trained according to this method may, for example, output a 3D modeled object based on a corresponding 3D modeled object in a dataset based on a latent vector representation from the latent space of the 3D generative model. The latent vector representation may be computed, for example, by sampling latent vectors from the latent space of the deep learning generative model and outputting a 3D modeled object based on such latent vectors. Alternatively, the latent representation may correspond to a 3D modeled object provided to the deep learning generative model and projected into the latent space (e.g., represented as a latent vector). The mapping distance specifies how the deep learning generative model outputs a 3D modeled object based on a corresponding 3D modeled object in a dataset by penalizing shape dissimilarity between each output 3D modeled object and the corresponding modeled object in the dataset. In an example, the mapping distance penalizes such shape dissimilarity by penalizing the distance between the latent vector representation of the corresponding modeled object in the dataset and the latent vector representation result of applying the mapping model to the latent vector representation of the corresponding 3D modeled object in the dataset. The mapping model can be any neural network, for example, two fully connected layers, and thus can be trained according to machine learning techniques. Thus, in the example, the minimum (found thanks to the mapping distance penalty) corresponds to the latent vector that is closest to the latent vector representation of the corresponding 3D modeled object in the dataset. Thus, the deep learning generative model outputs the 3D modeled object with the best shape consistency, i.e., the 3D modeled object that corresponds to the latent vector found thanks to the mapping distance penalty.Thus, the deep learning generative model outputs 3D modeled objects with improved feature and shape consistency with respect to the corresponding 3D modeled objects in the dataset (which may have been sampled or obtained from the provided 3D objects), resulting in more accurate output 3D modeled objects with respect to the corresponding 3D modeled objects in the dataset, with optimized features (thanks to feature loss) and improved shape consistency (thanks to mapping distance).
[0052] The training may include, for each 3D modeled object, calculating a feature score for the 3D modeled object, the calculation being performed by applying one or more of a deterministic function, a simulation-based engine, or a deep learning function to the 3D modeled object. Thus, the feature scores may input feature annotations into a dataset of the 3D modeled object.
[0053] By "deterministic function" is meant any function (including at least a sequence of calculations) that provides an explicit deterministic theoretical model, i.e., produces the same output from given starting conditions. This specifically excludes methods that are stochastic or comprise stochastic elements. A deterministic function evaluates a 3D object (e.g., provided as input to an explicit deterministic theoretical model) and outputs, at a minimum, a feature score that corresponds to a feature descriptor.
[0054] "Simulation engine" means any data structure representing a set of calculations that, as supported by the simulation engine (e.g., by virtue of a physics engine), at least partially evaluate the interactions of 3D objects under various interactions. A simulation-based engine takes 3D objects as input and outputs quantities related to the simulation output as feature scores. Thus, a simulation-based engine evaluates 3D objects within a relevant context, i.e., tracks the behavior of the 3D objects when subjected to various expected interactions, such as the expected response of a mechanical part or assembly of mechanical parts under the action of gravity.
[0055] "Deep learning function" means any data structure representing a sequence of computations, including learning feature scores based at least in part on a dataset, and may be based at least in part on one or more of a deterministic function or a simulation-based engine. Thus, a 3D object is fed into the deep learning function, which predicts feature scores that can be used as annotations for the 3D modeled object.
[0056] If the calculation is performed by a deep learning function, the deep learning function may have been trained on other datasets. The other datasets may include 3D objects each associated with a respective feature score. The respective feature scores may have been calculated using one or more of a deterministic function, a simulation-based engine, or a deep learning function. Thus, this method leverages the legacy of 3D objects generated using existing computer programs, which are characterized by highly heterogeneous compliance with functional requirements.
[0057] Thus, the training method may utilize at least a deterministic function, a simulation engine, and / or a deep learning function, or a combination of such computational tools to calculate the feature scores. In an example, the training may augment a dataset of 3D modeled objects with functional validity for each 3D modeled object in the dataset by using the calculated feature scores as annotations.
[0058] 4, an example of a computational tool used to calculate feature scores corresponding to particular categories of 3D modeled objects is provided. The dataset provided in S10 may include 3D modeled objects of any combination of one or more of these particular categories, and the one or more feature scores in S20 may include, for each particular category, one or more (e.g., all) feature descriptors shown in the table, optionally calculated as shown in the table.
[0059] The following documents are referenced below:
[10] Hongtao Wu & al. Is That a Chair? Imagining Affordances Using Simulations of an Articulated Human Body. ICRA 2020.
[11] Dule Shu & al. 3D Design Using Generative Adversarial Networks and Physics-Based Validation. 2019.
[12] Mihai Andries & al. Automatic Generation of Object Shapes With Desired Affordances Using Voxel grid Representation. 2019.
[13] Vladimir G. Kim & al. Shape2Pose: Human-Centric Shape Analysis. SIGGRAPH 2014.
[0060] In these examples, for the chair category, the computational tool comprises a deep learning function for calculating connectivity loss and a simulation engine for calculating physical stability and seating affordance descriptors. The one or more feature descriptors provided herein are merely examples. In this example, connectivity loss is calculated via topological priors, e.g., losses incorporated into the deep learning function as in [6], and is therefore referred to as a "topological loss." Furthermore, a simulation engine is used to calculate feature scores corresponding to the physical stability descriptor by applying gravity to the object. Furthermore, a simulation engine such as
[10] is used to calculate feature scores corresponding to the seating affordance descriptor. In these examples, for the airplane category, feature scores corresponding to the connectivity descriptor may be obtained using the topological priors described above. Furthermore, scores corresponding to the drag coefficient descriptor may be calculated via a simulation engine that performs computational fluid dynamics simulations as in
[11] . In these examples, for the pushcart vehicle category, a simulation engine such as
[12] may be used to further calculate feature scores corresponding to contain affordances. In these examples, a further category of vehicle may include bicycle objects, and in particular, compute feature scores corresponding to human support and / or pedal affordances. In this example, a feature energy model such as that used in
[13] may be incorporated into the simulation engine to provide an affordance model that can be used to compute such descriptors by simulating the application of forces to the bicycle seat, corresponding to the forces applied by a human model, as well as simulating the bicycle dynamics under pedaling effects.
[0061] One or more feature descriptors are now described.
[0062] The one or more feature descriptors may include a connectivity descriptor. A "connectivity descriptor" refers to any variable (one-dimensional or multi-dimensional) that represents the number of connected elements of a 3D modeled object. Therefore, the training method may focus on penalizing isolated elements of the 3D model. As a result, a deep learning generative model trained according to the training method creates a 3D modeled object composed of a single connected component. Therefore, because the resulting object is a single 3D modeled object without isolated components, the resulting 3D object is more accurate, a feature particularly required for the design of a mechanical part or assembly of mechanical parts.
[0063] The one or more functional descriptors may include one or more geometric descriptors and / or one or more affordances. A "geometric descriptor" refers to any variable (one-dimensional or multi-dimensional) that describes the spatial structure of a 3D object. An "affordance" refers to any variable (one-dimensional or multi-dimensional) that describes the interaction of the object within its intended context (i.e., any type of relationship that provides clues about the use of the 3D object within a particular context). Thus, the training method may focus on penalizing geometric aberrations. As a result, a deep learning generative model trained according to the training method provides geometrically stable 3D modeled objects, i.e., with more accurate spatial structure and / or better functionality regarding the interaction of the 3D modeled objects within their intended context.
[0064] The one or more geometric descriptors may include one or more of a physical stability descriptor and / or a durability descriptor. For a mechanical part or assembly of mechanical parts, the physical stability descriptor may represent the stability of the mechanical part or assembly of mechanical parts, e.g., the ability of the mechanical part or assembly of mechanical parts to maintain equilibrium in a spatial position under the application of gravity alone. Thus, the training method may focus on penalizing deviations of mechanical parts from their initial positions under the application of gravity. Thus, a deep learning generative model trained according to the training method outputs a 3D model object that maintains mechanical stability under the action of gravity, as expected in the context of a 3D object representing a mechanical part or assembly of mechanical parts.
[0065] In an example, a physical stability descriptor for a 3D object may be calculated via a deterministic function, a simulation engine, or a deep learning function. In such an example, the descriptor may correspond to the position of the center of gravity of a mechanical part or assembly of mechanical parts, and the response of the 3D object under the application of gravity is recorded over a time interval at least at those positions. Thus, a performance score measuring the degree of non-respect of the stability descriptor is the difference between the initial and final spatial positions of the center of gravity over the time interval. The physical stability descriptor may further be used to define a performance loss, such as:
[0066]
number
[0067] where:
[0068]
number
[0069] is the time from time i0 to time i frepresents the center of gravity position p ranging from . The definition of the time interval may be defined over, for example, a discrete interval, a continuous interval, or a discretization of the time interval. Such a loss penalizes the movement of the center of gravity under the action of gravity over the time interval.
[0070] The durability descriptor, for a mechanical part or assembly of mechanical parts, represents the ability of the mechanical part or assembly of mechanical parts to withstand the application of gravity and external mechanical forces. The external mechanical forces may correspond to forces applied at random locations, perturbing the positions by multiplying the mass and gravity at each of the random locations. The external mechanical forces may further correspond to forces characterized by direct contact with other mechanical objects, including, for example, any type of friction between two objects. Thus, the training method may focus on penalizing deviations in the spatial position of the 3D modeled object due to perturbations. As a result, a deep learning generative model trained according to the training method may output a 3D modeled object that provides a more accurate representation of the mechanical part or assembly of mechanical parts subjected to the application of gravity and external mechanical forces, for example, representing stress, vibration forces, or similar external mechanical perturbations.
[0071] In an example, the response of a 3D object may be calculated via a deterministic function, a simulation engine, or a deep learning function. In such an example, the durability descriptor may be the center of gravity location of a mechanical part or assembly of mechanical parts, and further, the initial and final spatial locations of the object's center of gravity are recorded to evaluate the durability of the 3D modeling. Thus, the durability descriptor is used to define a functional loss term that penalizes deviations from the spatial location of the center of gravity. The durability descriptor may be used to define the functional loss term as follows:
[0072]
number
[0073] where α is the annealing coefficient and the center of gravity is
[0074]
number
[0075] , k is a user-defined weight. The loss therefore penalizes objects that fail to remain stationary under perturbations. Objects that remain stationary under perturbations are a desirable property in the context of mechanical parts or assemblies of mechanical parts.
[0076] The one or more affordances may include one or more of a support affordance descriptor, a drag coefficient descriptor, a containment affordance descriptor, a holding affordance descriptor, and a suspension affordance descriptor. The support affordance descriptor, for a mechanical part or assembly of mechanical parts, represents the ability of the mechanical part or assembly of mechanical parts to withstand the application of only external mechanical forces. Such a descriptor may be a position (e.g., top position) of the mechanical part or assembly of mechanical parts, which may be recorded, for example, by a simulation engine or an explicit theoretical model, in which one or more forces are applied to the top of the mechanical part or assembly of mechanical parts. The drag coefficient descriptor, for a mechanical part or assembly of mechanical parts, may represent the effect of a fluid environment on the mechanical part or assembly of mechanical parts. The descriptor may be a drag coefficient, i.e., a dimensionless quantity used to represent the drag or resistance of a 3D modeled object in a fluid environment, as commonly known in the field of fluid dynamics. Such a descriptor may be calculated, for example, via a simulation engine performing a computational fluid dynamics simulation. A containment affordance descriptor may represent, for a mechanical part or assembly of mechanical parts, the response of the mechanical part or assembly of mechanical parts while containing other objects within the interior volume of the mechanical part. Such descriptors may be calculated, for example, via a simulation engine. A hold affordance descriptor may represent, for a mechanical part or assembly of mechanical parts, the ability of the mechanical part or assembly of mechanical parts to support other objects via mechanical connections. Such descriptors may be defined at locations where the mechanical connections are located. A suspension affordance descriptor may represent, for a mechanical part or assembly of mechanical parts, the ability of the mechanical part or assembly of mechanical parts to be supported via mechanical connections.
[0077] Each 3D modeled object in the dataset may represent furniture, a motorized vehicle, a non-motorized vehicle, or a tool. Thus, a deep learning generative model trained according to this method provides a more accurate representation of the specified class of objects in the dataset. The training may use any loss that penalizes combinations of feature scores, for example, by penalizing one or more combinations of a connectivity descriptor, one or more geometric descriptors, or one or more affordances. A furniture class may have as its functional descriptors at least a connectivity descriptor, a physical stability descriptor, and either a seating affordance or an object support affordance. A motorized vehicle class may have as its functional descriptors at least a connectivity descriptor and a drag coefficient descriptor. A tool class may have as its functional descriptors at least a connectivity descriptor, a durability descriptor, and either a holding affordance or an object support affordance.
[0078] The training method can combine one or more feature scores for each descriptor according to the intended context of the output 3D modeled object. See FIG. 4 for an example of calculating feature scores for object categories of chair, table, airplane, bed, push cart, and bicycle. Thus, as shown by the example, the training method considers the interdependence of the geometric, structural, topological, physical, and functional features of the 3D modeled object, ensuring that the output 3D modeled object is accurate, so that the generated content achieves its purpose and can be physically manufactured and brought into the real world.
[0079] An example will now be described with reference to Figures 5 to 7.
[0080] FIG. 5 shows an example of synthesis of multiple 3D modeled objects in the specific category of chairs by a deep learning generative model trained according to the training method.
[0081] Figure 6 illustrates the synthesis of a 3D modeled object of a table category, where a deep learning generative model consisting of a variational autoencoder and trained according to our training method. Two latent vector representations are sampled and output, as shown on the far left and right of Figure 6. The 3D modeled object in the center corresponds to the interpolation of the two sampled latent vector representations. This example demonstrates how a deep learning generative model trained according to our method can leverage existing 3D objects, which are automatically synthesized using state-of-the-art generative networks (in this example, in the form of a variational autoencoder), to output a functionally valid 3D modeled object.
[0082] 7 shows an example of the incorporation of a deep learning generative model within a design process, particularly in a CAD environment. The trained deep learning generative model maps (evaluates and updates) input modeled objects to output 3D modeled objects, on which subsequent design (edits) can be performed. An example will now be described with reference to Figures 8 to 16.
[0083] In these examples, each 3D modeled object in the dataset represents a piece of furniture, specifically a chair, i.e., the training and usage methods apply to the category "chairs". The 3D data representation consists of structured point cloud objects, i.e., each 3D modeled object is a point cloud.
[0084] The following example relates more to computing feature descriptors and scores, connectivity, stability, durability, and seating affordance descriptors for a specific category of chairs, a class of furniture objects. The feature descriptors are used to generate a structured point cloud object representing the category.
[0085] Connectivity Descriptor: Within the category of chairs, connectivity encourages the absence of floating segments of objects. Referring to Figure 8, (a) shows an object that violates the connectivity descriptor. Given the category, having a connected object corresponds to having one connected component.
[0086] In these examples, to evaluate the connectivity of point clouds, the method of [6] is used to calculate the connected and separated components of a 3D modeled object using 0d persistent homology. According to the method of [6], when a topological feature (connected component) appears (birth time b i ), or when it disappears (death time d i ) can be calculated.
[0087] In these examples, the connectivity loss is defined as:
[0088]
number
[0089] This loss function is summed over a lifetime starting at i=1, so the loss penalizes isolated components.
[0090] Stability Descriptor: 3D objects are input into a simulation engine [8] to simulate the stability of objects in the chair category. The simulation engine evaluates the stability of 3D objects by assessing whether they remain stationary when subjected to gravity after being placed on a plane with a common orientation. Thus, the simulation is performed using a 2.510 -4 In the simulation step of seconds, from i0=0 to i 2s Record the position of the center of gravity, pi, when subjected to gravity for the range i = 2.5 seconds. The corresponding functional loss is:
[0091]
number
[0092] Referring to Figure 8, (b) shows an object that violates the stability descriptor. Figure 9 shows an example simulation showing the initial and final frames of the simulation for the chair category. The top shows the initial frame and the bottom shows the final frame, with the initial time being 0 and the final time being 2.5 seconds. The top and bottom frames in the center of Figure 9 show an unstable object (high f stability ) and the top and bottom of the side frames are stable objects (small f stability ) is shown.
[0093] Durability Descriptor: Ensures that the chair remains stable when subjected to small perturbations. Referring to Figure 8, (c) shows an object in the chair category that violates the durability descriptor. In this example, the perturbing force tends to rotate the object in the xy (horizontal) plane. The force is applied at random positions and has 10 different norms: k * 0.01 times the gravity norm of the object's mass, for k ranging from 1 to M. The object's behavior is simulated for each perturbation, and the initial and final positions of the object's center of mass are recorded, as described in the paragraph above, respectively.
[0094]
number
[0095] and
[0096]
number
[0097] In this example, the loss of functionality is:
[0098]
number
[0099] where α is an annealing coefficient between 0 and 1 that ensures that objects in the chair category are penalized more for their inability to remain stationary under smaller perturbations. In the current example, α=0.9 and M=10.
[0100] The following example shows how feature descriptors can be further combined: stability and f durability are combined here as follows:
[0101]
number
[0102] Figure 10 shows the trajectories recorded for different values of k for the two chairs. It can be interpreted from the trajectories at the bottom of Figure 10 that the object on the left has a lower f than the object on the right. s+d (i.e., it has better functionality). The trajectory on the bottom left shows that the corresponding object on the top left remains stationary when subjected to gravity for k = 0 (blue horizontal curve), and recovers its initial position in response to gravity and various small perturbations for k > 1. In contrast, the chair on the right fails to recover its initial position, as is evident from the trajectory on the bottom right of Figure 10.
[0103] In a further example, all three scores may be further combined as follows:
[0104]
number
[0105] In the following example, the loss is calculated using three feature scores f physical , and furthermore, a feature score f that measures the degree of non-respect for each locus affordance descriptor. affordanceReference is made to FIG. 11, which illustrates the calculation of feature scores. In this example, the two objects on the top initial frame both have a score of
[0106]
number
[0107] , and thus maintains a stationary state under gravity. However, the chair on the right in Figure 11 is clearly non-functional due to the absence of a seat portion and should show a higher functionality score compared to the chair on the left. Therefore, a simulation engine such as that in [8] is used to calculate the functionality score of the sitting affordance descriptors, i.e., to evaluate how well the object satisfies the sitting affordance requirements. In the simulation engine, a human agent is provided as shown in Figure 11 and is free to drop onto the object. The simulation engine evaluates the response of the chair category in the context of the human agent's interaction. The human agent is an articulated human body with 18 joints and composed of 9 links. In this example, the arms and legs are not important in defining a sitting configuration, so they are trimmed. Appropriate limits, friction, and damping for each joint are set to avoid configurations that are physiologically impossible for a typical human. The agent's pelvis is positioned on a horizontal plane 15 cm above the bounding box aligned with the object's current axis. Each drop is a sitting trial. The agent's resulting configuration C res is recorded, which corresponds to a 24-dimensional vector including the center of gravity position (3 dimensions), direction (3 dimensions), and joint angles (18 dimensions). As shown by the initial frame in Figure 11, the key C corresponding to the target position key The locus configuration is provided. In this example, the locus affordance score f affordance is as follows:
[0108]
number
[0109] The following example shows the implementation of a deep learning generative model trained according to the training method.
[0110] The implemented deep learning model takes a 3D object as input and outputs a feature score f physical and f affordance To build the model, we use a 3D generative neural network such as that described in [9, 10], where N=5.10. 4 A database of chairs is generated. Specifically, given an object category, a pre-trained 3D generative neural network from the prior art is used to sample new content of this category, as directed by each of these publications. Typically, in the case of a GAN-based 3D generative neural network, such as [8], trained to learn the data distribution of the target object category, many sampled vectors are mapped from the learned distribution to a voxel grid as outputs using a model generator. Each of these outputs constitutes a new instance. Each chair object O i For the function score f i is calculated as above, which gives us the new labeled dataset
[0111]
number
[0112] The feature predictor is created by this estimated vector f i ' and ground truth feature scores f i The training is to reduce the distance between the two. The distance is determined by a loss function such as Euclidean distance.
[0113] In this example, the feature score is defined as:
[0114]
number
[0115] All scores are normalized between 0 and 1. Figure 12 shows the generated object Oi with the corresponding f i are shown in descending order.
[0116] To train a deep learning generative model,
[0117]
number
[0118] is mapped to a point cloud, using the PointNet architecture as in [9]. In this example, the training loss for the feature predictor is:
[0119]
number
[0120] The following example illustrates the use of a deep learning generative model. In the following, we refer to the portion of the deep learning generative model trained according to the training method that outputs feature scores for a corresponding output 3D modeled object as the "feature predictor." In this example, the deep learning generative model consists of a 3D generative neural network followed by a feature predictor. In this example, the 3D generative neural network corresponds to a PQ-Net decoder [8].
[0121] The PQ-Net decoder receives latent vectors from the learned latent space of the object and maps them to 3D objects one at a time to create a sequential assembly. The decoder performs shape generation by training a generator using a GAN strategy as described in [8]. The GAN generator maps random vectors sampled from a standard Gaussian distribution N(0,1) to latent vectors in the object latent space of the object, from which the sequential decoder generates a new object.
[0122] First use.
[0123] The first usage method is an implementation example of a deep learning generative model. The deep learning generative model is composed of a 3D generative neural network, which is a generative adversarial network, and the other term includes an adversarial loss. In this example, the 3D generative neural network is trained to map latent vectors to 3D objects while ensuring that the output content has a low feature score. Figure 13 provides an overview of the first usage method. In Figure 13, the feature predictor is pre-trained, and the deep learning generative model is composed of a 3D generative network, which can be trained according to the training method. The deep learning generative model architecture consists of a 3D generative neural network (PQ-Net decoder) followed by a feature predictor. During the training process, the weights of the 3D generative neural network are trained, and the weights of the feature predictor are pre-trained and frozen as described above.
[0124] At each training iteration, the latent vector Z in is sampled from the latent space and fed to the generator, and the latter is Z in 3D object (O i ) to map to a 3D object (O i ) is its feature score f i The network then calculates the feature scores fi We generate a geometrically plausible 3D object (O i ), the model is therefore equipped with functional reasoning. The latent representation contains both the geometric (here including structure) and functional dimensions of the 3D object. The training loss is the feature score L f =f i and a term that penalizes adversarial losses.
[0125]
number
[0126] A deep learning generative model trained according to this method synthesizes new objects, as shown in Figure 6.
[0127] Second usage.
[0128] The second usage method shows an implementation in which the deep learning generative model consists of a mapping model followed by a 3D generative neural network, where the 3D generative neural network is pre-trained, other terms include the mapping distance, and the 3D generative neural network is optionally a variational autoencoder. An overview of the deep learning generative model is shown in Figure 14. In Figure 14, the 3D generative neural network is pre-trained, and a feature predictor is included in the deep learning generative model and is also pre-trained. The mapping model shown in Figure 4 is trained alone according to this method, and the remaining boxes correspond to the pre-trained and frozen network. This usage method optimizes the functionality of the output 3D object, as described below. In this example, the deep learning generative model consists of a pre-trained 3D generative neural network (PQ-Net decoder), the pre-trained feature predictor described above, and a mapping model consisting of a neural network with two fully connected layers trained according to the training method. In this usage method, the deep learning generative model generates a latent vector representation Z of each 3D object. in This method is used to optimize Z, which minimizes the feature score. in The latent vector Z closest to out Therefore, the provided 3D object is transformed into a latent vector representation Z with improved functionality. out (i.e., objects with low feature scores). in From Z out The mapping to is achieved using a mapping model, which in this implementation consists of a neural network with two fully connected layers. An overview of the model is illustrated in Figure 14.
[0129] Therefore, the mapping distance is used to train a mapping model, which involves minimizing a loss that includes the mapping distance, penalizing the feature scores.
[0130]
number
[0131] The weights of the remaining models (generative models and feature predictors) are frozen.
[0132] The network then computes the optimal Z generated by the generative model. out and returns the corresponding 3D object. Figure 15 shows an example of optimization, showing a pair of objects in each corner. Each pair shows the sampled object before optimization on the left and the optimized output object on the right. For example, object 4 in the bottom right corner shows a sample consisting of a chair with three legs and a defective fourth leg. After optimization, the output diagram shows the fourth leg without a defect, and therefore has improved functionality. The results of the numerical optimization are given by Figure 16, which shows the improvement in functionality of the output 3D modeled object after optimization.
Claims
1. 1. A computer-implemented method for training a deep learning generative model that outputs 3D modeled objects, each representing a mechanical part or an assembly of mechanical parts, comprising: - providing a dataset of 3D modeled objects, each 3D modeled object representing a mechanical part or an assembly of mechanical parts; training the deep learning generative model based on the dataset, wherein the training generates a loss (L Train , L map ), and minimizing the loss (L Train , L map ) represents the respective output 3D modeled object (O i ) for each 3D modeled object (O i ) one or more functional scores (f i ) and calculate the score of each function (f i ) measuring the degree of non-respect of each functional descriptor among the one or more functional descriptors by the mechanical part or assembly of the mechanical parts; Including, The training includes, for each 3D modeled object (Oi), calculating a feature score (f i ) for the 3D modeled object (Oi), the calculation including assigning to the 3D modeled object (Oi): a deterministic function (φ det (·)), Simulation-based engines, or Deep learning function (φ DL (・)) The method is carried out by applying one or more of the following:
2. The loss (L Train , L map ) is the output 3D modeled object (O i ) for each of the respective 3D modeled objects (O) associated with the data set i ) shape mismatch (||Z out -Z in || 2 , L VAE , L GAN 2. The method of claim 1 further comprising another term that penalizes .
3. The other terms are: Each of the 3D modeled objects (O i ) and the corresponding ground truth 3D modeled object of the dataset, VAE ), The adversarial loss (L GAN ),or, Each of the 3D modeled objects (O i ) and the corresponding modeled object in the dataset. out -Z in || 2 ) The method of claim 2 , comprising:
4. The method of claim 3 , wherein the deep learning generative model comprises a 3D generative neural network.
5. The method of claim 4 , wherein the 3D generative neural network comprises a variational autoencoder or a generative adversarial network.
6. The deep learning generative model is the 3D generative neural network comprising a variational autoencoder, and the other terms comprising a reconstruction loss and a variational loss; The 3D generative neural network includes a generative adversarial network, and the other term is an adversarial loss (L GAN ) or A mapping model followed by the 3D generative neural network, wherein the 3D generative neural network is pre-trained and the other term is the mapping distance (||Z out -Z in || 2 ), wherein the 3D generative neural network optionally comprises a variational autoencoder or a generative adversarial network. The method of claim 4 , comprising one of:
7. The calculation is performed by using the deep learning function (φ DL (·)), and the deep learning function (φ DL (·)) is trained based on another dataset, the other dataset including 3D objects each associated with a respective feature score, the respective feature scores being: A deterministic function (φ det (・)), Simulation-based engines, or Deep learning function (φ DL (・) The method of claim 1 , wherein the calculation is performed by using one or more of:
8. The method of claim 1 , wherein the one or more feature descriptors include a connectivity descriptor.
9. The one or more feature descriptors: one or more geometric descriptors, and / or One or more affordances 9. The method of claim 1, comprising:
10. The one or more geometric descriptors are: a physical stability descriptor, for a mechanical part or an assembly of mechanical parts, that describes the stability of said mechanical part or said assembly of mechanical parts under the application of gravity alone; and / or A durability descriptor for a mechanical part or an assembly of mechanical parts, the durability descriptor describing the ability of said mechanical part or said assembly of mechanical parts to withstand the application of gravity and external mechanical forces. and / or The one or more affordances are: a support affordance descriptor, for a mechanical part or an assembly of mechanical parts, that describes the ability of said mechanical part or said assembly of mechanical parts to withstand the application of external mechanical forces only; and / or a drag coefficient descriptor, for a mechanical component or an assembly of mechanical components, that describes the effect of a fluid environment on said mechanical component or said assembly of mechanical components; a containment affordance descriptor, for a mechanical part or an assembly of mechanical parts, that describes the ability of the mechanical part or the assembly of mechanical parts to contain other objects in an interior volume of the mechanical part or the assembly of mechanical parts; a hold affordance descriptor, for a mechanical part or an assembly of mechanical parts, that describes the ability of said mechanical part or said assembly of mechanical parts to support other objects via a mechanical connection; and / or A hanging affordance descriptor, for a mechanical part or an assembly of mechanical parts, that describes the ability of the mechanical part or the assembly of mechanical parts to be supported via a mechanical connection.
10. The method of claim 9, comprising:
11. Each 3D modeled object in the dataset comprises: A piece of furniture, Electric vehicles, Non-motorized vehicles, or tool The method of claim 10, wherein 12. A computer program comprising instructions for carrying out the method according to any one of claims 1 to 11.
13. A device including a data storage medium having recorded thereon a computer program according to claim 12.
14. The device of claim 13 , further comprising a processor coupled to the data storage medium.
Citation Information
Patent Citations
Set of neural networks
JP2020115337A
3D Reconstruction Method Based on Deep Learning
US20200294309A1