Basic learning of transformation

A neural network-based method for inferring deformation bases in 3D modeled objects addresses inefficiencies in existing deformation methods by enabling efficient and generalized deformations across categories with reduced computational resources and dataset requirements.

JP7730264B2Active Publication Date: 2025-08-27DASSAULT SYSTEMES SA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2021011110
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-01-30
Filing Date
2021-01-27
Publication Date
2025-08-27
Estimated Expiration
2041-01-27

AI Technical Summary

Technical Problem

Existing systems lack efficient and generalized methods for deforming 3D modeled objects, particularly in terms of computation time and resource usage, and require specific category labeling and large datasets for effective deformation.

Method used

A neural network is trained to infer a deformation basis for 3D modeled objects, allowing for efficient deformation calculation by linear combinations without requiring target deformations, and can generalize across different categories even with smaller datasets.

Benefits of technology

The method enables efficient and realistic deformations of 3D modeled objects, reducing computation time and resource usage, and improves generalization ability by correlating deformations across categories, even with smaller datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007730264000035
    Figure 0007730264000035
  • Figure 0007730264000036
    Figure 0007730264000036
  • Figure 0007730264000037
    Figure 0007730264000037
Patent Text Reader

Abstract

To provide a machine learning method of deforming a 3D model object, a system therefor, and a program therefor.SOLUTION: A method by a deep feed-forward neural network has a step of providing data set for 3D model object and a step of learning a neural network which deducts a deformation basis of the input 3D model object. The neural network has an encoder for receiving the 3D model object as an input and outputting a potential vector representing the input 3D model object. The potential vector output from the encoder is received as an input, and the deformation basis of the 3D model object represented by the potential vector is output.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the field of computer programs and systems, and more particularly to machine learning methods, systems and programs for deforming 3D modeled objects. [Background technology]

[0002] Many systems and programs are available on the market for designing, engineering, and manufacturing objects. CAD stands for Computer-Aided Design and refers to software solutions for designing objects. CAE stands for Computer-Aided Engineering and refers to software solutions for, for example, simulating the physical behavior of future products. CAM stands for Computer-Aided Manufacturing and refers to software solutions for, for example, defining manufacturing processes and operations. In such computer-aided design systems, the graphical user interface plays a key role in the efficiency of the technology. These technologies may be integrated within a Product Lifecycle Management (PLM) system. PLM is a business strategy that allows companies to share product data, apply common processes, and leverage corporate knowledge to support product development from conception to the end of product life, going beyond the concept of the extended enterprise. Dassault Systèmes' PLM solutions (under the trademarks CATIA, ENOVIA and DELMIA) include an Engineering Hub to organize product engineering knowledge, a Manufacturing Hub to manage manufacturing engineering knowledge, and an Enterprise Hub to enable enterprise integration and connectivity to both the Engineering and Manufacturing Hubs. The system provides an open object model that links products, processes and resources, enabling dynamic knowledge-based product creation and decision support, facilitating the optimization of product definition, manufacturing preparation, production and service.

[0003] In this and other contexts, deformation of 3D modeled objects has become widely important.

[0004] The following papers are relevant to this field and are cited below: [1] W. Wang, D. Ceylan, R. Mech, and U. Neumann. 3Dn: 3D deformation network. In Conference on Computer Vision and Pattern Regognition (CVPR), 2019. [2] T. Groueix, M. Fisher, V. G. Kim, B. Russell, and M. Aubry. AtlasNet: A Papier-Mache Approach to Learning 3D Surface Generation. In Conference on Computer Vision and Pattern Regognition (CVPR), 2018. [3] Reconstructing a 3D Modeled Object. Patent US9978177B2, granted in 2018. Eloi Mehr and Vincent Guitteny. [4] D. Jack, J. K. Pontes, S. Sridharan, C. Fookes, S. Shirazi, F. Maire, and A. Eriksson. Learning free-form deformations for 3D object reconstruction. In Asian Conference on Computer Vision (ACCV), 2018. [5] I. Kokkinos, and A. Yuille. Unsupervised Learning of Object Deformation Models. In International Conference on Computer Vision (ICCV), 2007. [6] J. Mairal, F. Bach, and J. Ponce. Sparse Modeling for Image and Vision Processing. New Foundations and Trends, 2014. [7] V. Blanz and T. Vetter. A morphable model for the synthesis of 3D faces. In SIGGRAPH, 1999. [8] C. Qi, H. Su, K. Mo, and L. Guibas. Pointnet: Deep learning on point sets for 3D classification and segmentation. In Conference on Computer Vision and Pattern Regognition (CVPR), 2017. [9] R. Hanocka, A. Hertz, N. Fish, R. Giryes, S. Fleishman, and D. Cohen-Or. Meshcnn: A network with an edge. In SIGGRAPH, 2019.

[10] Eloi Mehr, Ariane Jourdan, Nicolas Thome, Matthieu Cord, and Vincent Guitteny. DiscoNet: Shapes Learning on Disconnected Manifolds for 3D Editing. In ICCV 2019.

[0005] However, there is still a need for improved solutions for deforming 3D modeled objects. Summary of the Invention

[0006] Accordingly, there is provided a computer-implemented method for machine learning, comprising providing a dataset of 3D modeled objects. The method further comprises training a neural network, the neural network configured to infer a deformation basis for the input 3D modeled objects. The method may be referred to as a "training method."

[0007] This constitutes an improved solution for deforming 3D modeled objects.

[0008] It is noteworthy that the neural network trained by the learning means is configured to take a 3D modeled object as input and output a deformation basis for the 3D modeled object. The deformation basis is composed of vectors that form a deformation of the input 3D modeled object. The vectors of the deformation basis are linearly combined, and each linear combination results in a different deformation of the 3D modeled object, thereby obtaining many deformations of the 3D modeled object. In this way, the learning method increases the number of possible deformations of the 3D modeled object. Furthermore, the learning method allows the deformation basis to be calculated once by a single application of the trained neural network, so that only linear combinations need to be performed thereafter. This allows for efficient deformation calculation, particularly in terms of computation time and resource usage. It is also noteworthy that the neural network does not require a target deformation of the 3D modeled object to obtain the deformation basis; the neural network alone can obtain a basis of linearly combined deformation vectors to arrive at the target deformation of the input 3D modeled object. In this way, the learning method has a large generalization ability and maximizes the expressive power of the input 3D modeled object.

[0009] The neural network may infer a local deformation basis consisting of local deformation vectors, i.e., vectors each representing a local (e.g., small) deformation of an input 3D modeled object, which allows the 3D modeled object to deform realistically in a local neighborhood of the 3D modeled object (i.e., in the vicinity of the 3D modeled object). In other words, the neural network outputs a satisfactory basis in the sense that it allows the input object to deform towards nearby objects in a physically acceptable way.

[0010] Furthermore, the neural network is trained on the 3D modeled objects of the dataset, at least a majority (e.g., overall) of which are plausible 3D modeled objects. This allows the neural network to obtain a basis of plausible vectors of deformations of the input 3D modeled objects. As mentioned above, these vectors may be linearly combined to obtain plausible deformations, for example, in a local neighborhood of the input 3D modeled object.

[0011] Furthermore, for a neural network to be trained to infer the basis of plausible deformation vectors for input 3D modeled objects, the provided dataset does not need to be made up of objects of the same category (e.g., all chairs). Specifically, if the provided dataset is made up of, or substantially made up of, a large dataset of 3D modeled objects of the same category (e.g., all chairs), the neural network is allowed to infer physically realistic deformation bases for objects of this category (e.g., chairs). However, the neural network may do so even if the 3D modeled objects of the provided dataset form different categories (e.g., chairs and benches), and even if these categories are not represented by a large number of 3D modeled objects in the provided dataset. Notably, training does not require that the objects in the provided dataset be annotated with labels (if any) indicating their respective categories. Nor does training require the use of template modeling objects, such as the average shape of a category of objects. Furthermore, the provided dataset does not need to be clustered according to object categories. This further improves the generalization ability of the neural network.

[0012] Nevertheless, if the provided dataset consists of two or more categories of 3D modeled objects that form a manifold (e.g., chairs and benches), the generalization ability of the neural network is further improved. In fact, a manifold may be formed by different categories of objects that have at least two-to-two connectivity between them, i.e., a certain amount of shared features / characteristics (e.g., as chairs and benches do). Therefore, a single vector of deformation bases may be a meaningful and plausible deformation for many object categories in the manifold (e.g., the seats of chairs and benches may be similarly deformed). In other words, if the provided dataset forms a manifold of 3D modeled objects (e.g., chairs and benches), the neural network is trained in a way that correlates between deformations of different categories of objects. This allows the neural network to learn to infer physically realistic / acceptable deformation criteria even if the provided dataset contains only a small number of modeled objects of different categories that form the manifold. This improves the generalization ability of the neural network.

[0013] Generally speaking, the generalization power that learning gives to a neural network is such that, compared to other deep learning frameworks, a training dataset (i.e., the provided dataset) with a smaller number of 3D modeled objects per category is required without affecting the neural network's ability to infer the basis of deformations. Specifically, the number of objects per category is smaller because the neural network can correlate learned deformations in different categories. However, if the provided dataset consists of only one category of two very different categories of objects, a larger number of objects may be required.

[0014] The deformations obtained using the deformation basis output by the neural network may be used for 3D modeling, such as in 3D editing applications. The deformation basis inferred by the neural network may also be used for other applications, such as shape matching and nonlinear deformation inference, as described below. Furthermore, because neural networks provide linear deformations of input 3D modeled objects, they are efficient for use in real-time applications.

[0015] The learning method may consist of one or more of the following: the neural network, an encoder configured to receive a 3D modeled object as input and to output a latent vector representing the input 3D modeled object; The deep feedforward neural network is configured to take the latent vectors output from the encoder as input and output a deformation basis for the 3D modeled object represented by the latent vectors. the training includes, for at least a portion of the dataset, for each 3D modeled object in the at least a portion of the dataset, for each candidate deformation basis having vectors, minimizing a loss that penalizes a distance between a deformation of the 3D modeled object and a deformation of another 3D modeled object by a linear combination of vectors. The training is performed mini-batch by mini-batch, and includes minimizing the loss for each mini-batch. the learning includes selecting the further 3D modeled object from among the at least some 3D modeled objects of the dataset based on a distance from the at least some 3D modeled object of the dataset. the other 3D modeled object is the 3D modeled object that is closest to the 3D modeled object of the at least part of the dataset among the 3D modeled objects of the at least part of the dataset. The learning is performed for each mini-batch, and includes minimizing a loss for each mini-batch and selecting the closest 3D modeled object from among the 3D modeled objects of the mini-batch. The loss penalizes the minimum value of the distance between the transformation of the 3D modeled object by a linear combination of the vector and the other 3D modeled object. The loss is of the following type:

number

number

[0016] Further provided are neural networks that can be trained according to the training method, for example, neural networks that have been trained (ie, learned) directly by the training method.

[0017] A method for using a neural network is also provided.

[0018] The method of use may consist of providing a neural network (e.g., by performing a training method) and applying the neural network to one or more 3D modeled objects (i.e., using the neural network to infer one or more deformation bases for input 3D modeled objects, respectively). The method of use may form an application of the training method. The application may be depth frame reconstruction, shape matching, or 3D editing. The method of use may be integrated into the training method, for example, as a step performed after training of the neural network, in which case the training method and the method of use form a computer-implemented method for inferring deformations of 3D modeled objects.

[0019] There is further provided a computer program comprising instructions which, when executed on a computer, cause the computer to carry out the method of learning and / or the method of use.

[0020] Further provided is a computer readable data storage medium having recorded thereon a computer program and / or a neural network.

[0021] Additionally, a computer is provided that comprises a processor coupled to a memory, the memory having a computer program and / or a neural network stored therein. [Brief explanation of the drawings]

[0022] [Figure 1] 1 shows an example of a computer. [Figure 2] Here's how. [Figure 3] Here's how. [Figure 4] Here's how. DETAILED DESCRIPTION OF THE INVENTION

[0023] The learning and use methods are computer-based.

[0024] This means that the method steps (or substantially all steps) are performed by at least one computer, or similarly any system. Thus, the method steps are performed by a computer, possibly fully automatically or semi-automatically. In examples, triggering of at least some of the method steps may be performed by user / computer interaction. The level of user / computer interaction required depends on the expected level of automation and can be balanced against the need to implement user wishes. In example embodiments, this level may be user-defined and / or predefined.

[0025] A typical example of a computer-implemented method is performing the method using a system adapted for this purpose. The system comprises a processor coupled to a memory and a graphical user interface (GUI), the memory having recorded thereon a computer program containing instructions for performing the method. The memory may also store a database. The memory may be any hardware adapted for such storage and may consist of several physically distinct parts (e.g., one for the program and one for the database).

[0026] Methods typically operate on modeled objects. A modeled object is an object defined by data stored in a database or other storage. By extension, the term "modeled object" refers to the data itself. Depending on the type of system, modeled objects are defined by different types of data. A system may actually be any combination of CAD, CAE, CAM, PDM, and / or PLM systems. In these different systems, modeled objects are defined by their corresponding data. Thus, one might speak of CAD objects, PLM objects, PDM objects, CAE objects, CAM objects, CAD data, PLM data, PDM data, CAM data, and CAE data. However, these systems are not exclusive, as modeled objects may be defined by data corresponding to any combination of these systems. Thus, a system may be both a CAD and a PLM system.

[0027] A CAD system additionally refers to a system, such as CATIA, adapted to design a modeled object based on at least a graphical representation of the modeled object. In this case, data defining the modeled object consists of data enabling the representation of the modeled object. A CAD system may, for example, provide a representation of a CAD modeled object using edges or lines, in certain cases with faces or surfaces. Lines, edges, or surfaces may be represented in various ways, for example, with non-uniform rational B-splines (NURBS). Specifically, a CAD file contains specifications from which geometry is generated, thereby generating a representation. The specifications of a modeled object may be stored in a single CAD file or in multiple CAD files. Typical sizes of files representing modeled objects in a CAD system are in the range of one megabyte per part. A modeled object may then typically be an assembly of several thousand parts.

[0028] In the context of CAD, a modeled object is typically a 3D modeled object, representing, for example, a part, an assembly of parts, or possibly a product such as an assembly of products. A "3D modeled object" means any object modeled by data that allows for a 3D representation. A 3D representation allows for viewing the part from any angle. For example, a 3D modeled object, when represented in 3D, can be manipulated or rotated around any of its axes or around any of the axes of the screen on which the representation is displayed. This notably excludes 2D icons that are not 3D modeled. Displaying 3D representations facilitates design (i.e., statistically improves the speed at which designers accomplish tasks). In industry, this translates to speeding up the manufacturing process, since product design is part of the manufacturing process.

[0029] A 3D modeled object may represent the shape of a product that will be manufactured in the real world after the virtual design is completed using, for example, a CAD software solution or CAD system, such as a (e.g., mechanical) part or assembly of parts (or equivalently, an assembly of parts, which may be viewed as a part itself from a method perspective, or the method may be applied to each part of the assembly independently), or more generally any rigid assembly (e.g., moving mechanism). CAD software solutions enable the design of products in a variety of unlimited industry sectors, including aerospace, architecture, construction, consumer goods, high-tech equipment, industrial equipment, transportation, marine, and / or offshore oil and gas production or transportation. Any 3D modeled object involved in the method can therefore represent an industrial product which may be any mechanical part, for example, a part of a land vehicle (including, for example, automobiles and light truck equipment, racing cars, motorcycles, trucks and motor equipment, trucks and buses, trains), a part of an air vehicle (including, for example, airframe equipment, aerospace equipment, propulsion equipment, defense products, aviation equipment, space equipment), a part of a naval vehicle (including, for example, naval equipment, commercial vessels, marine equipment, yachts and workboats, marine equipment), a general mechanical part (including, for example, industrial manufacturing machinery, heavy mobile machinery or equipment, installed equipment, industrial equipment products, fabricated metal products, tire manufacturing products), an electric machine or electronic part (including, for example, consumer electronics products, security and / or control and / or instrumentation products, computing and communications equipment, semiconductors, medical equipment and devices), a consumer product (including, for example, furniture, home and garden products, leisure products, fashion products, hard goods retail products, soft goods retail products), packaging (including, for example, food and beverage and tobacco, beauty and personal care, household product packaging).

[0030] FIG. 1 shows an example of a system, where the system is a client computer system, eg, a user's workstation.

[0031] The client computer of this example comprises a central processing unit (CPU) 1010 connected to an internal communication BUS 1000 and a random access memory (RAM) 1070 also connected to the BUS. The client computer further comprises a graphical processing unit (GPU) 1110 associated with a video random access memory 1100 connected to the BUS. The video RAM 1100 is also known in the art as a frame buffer. A mass storage controller 1020 manages access to mass memory devices such as a hard drive 1030. Mass memory devices suitable for embodying computer program instructions and data include, by way of example, semiconductor memory devices such as EPROMs, EEPROMs, and flash memory devices, magnetic disks such as internal hard disks and removable disks, magneto-optical disks, and all forms of non-volatile memory, including a CD-ROM disk 1040. Any of the foregoing may be supplemented by or incorporated into specially designed application-specific integrated circuits (ASICs). A network adapter 1050 manages access to a network 1060. The client computer may also include a haptic device 1090, such as a cursor control device, keyboard, or the like. A cursor control device is used in the client computer to allow a user to selectively position a cursor at any desired location on the display 1080. Furthermore, the cursor control device allows a user to select various commands and input control signals. The cursor control device includes a number of signal generating devices for inputting control signals to the system. Typically, the cursor control device may be a mouse, and the buttons on the mouse are used to generate the signals. Alternatively or additionally, the client computer system may have a sensitive pad and / or a sensitive screen.

[0032] A computer program may include computer-executable instructions, which include means for causing a system to perform a method. The program may be recordable on any data storage medium, including the system's memory. The program may be implemented, for example, in digital electronic circuitry, or computer hardware, firmware, software, or a combination thereof. The program may also be implemented as an apparatus, e.g., an article tangibly embodied in a machine-readable storage device for execution by a programmable processor. The method steps may be performed by a processor capable of executing a program of instructions to perform the functions of the method by operating on input data and generating output. Thus, the processor may be programmable and coupled to receive data and instructions from, and transmit data and instructions to, a data storage system, at least one input device, and at least one output device. The application program may be implemented in a high-level procedural or object-oriented programming language, or in assembly or machine language as desired. In either case, the language may be a compiled or interpreted language. The program may be a full installation program or an update program. In any case, applying the program on a system provides instructions for performing the method.

[0033] We will discuss providing datasets for 3D modeled objects. Prior to this discussion, we will discuss the data structures involved.

[0034] As used herein, any 3D modeled object may form a discrete geometric representation of a 3D shape, and may represent, for example, an object from the real world, such as a machine part as previously discussed. A discrete geometric representation, as used herein, is a data structure consisting of a discrete set of data. Each piece of data represents a respective geometric entity located in 3D space. Each geometric entity represents a respective position of the 3D shape (in other words, a respective piece of material that makes up the solid represented by the 3D shape). The collection (i.e., combination or juxtaposition) of the geometric entities as a whole represents the 3D shape. A discrete geometric representation, as used herein, may be composed of, for example, 100, 1000, or even higher than 10,000 pieces of data.

[0035] Any discrete geometric representation herein may be, for example, a 3D point cloud, where each geometric entity is a point. Any discrete geometric representation herein may alternatively be a 3D mesh, where each geometric entity is a mesh tile or face. Any 3D mesh herein may be regular or irregular (i.e., composed of faces of the same type or not). Any 3D mesh herein may be a polygonal mesh, for example, a triangular mesh. Any 3D mesh herein may be obtained from a 3D point cloud, for example, by triangulating the 3D point cloud (e.g., Delaunay triangulation). Any 3D point cloud herein may be determined from physical measurements on a real object, for example, within a 3D reconstruction process. The 3D reconstruction process may comprise providing the real object, providing one or more physical sensors, each configured to acquire a respective physical signal, and acquiring one or more respective physical signals by operating the one or more physical sensors on the real object (i.e., scanning the real object with the respective sensors). The 3D reconstruction may then automatically determine a 3D point cloud and / or a 3D mesh based on the measurements according to any known technique. The one or more sensors may consist of multiple (e.g., RGB, and / or image or video) cameras, and the determination may consist of structure-motion analysis. The one or more sensors may alternatively or additionally consist of one or more depth sensors (e.g., RGB depth cameras), and the determination may consist of 3D reconstruction from the depth data. The one or more depth sensors may consist of, for example, lasers (such as lidar) or ultrasonic emitters / receivers.

[0036] Any 3D point cloud or 3D mesh described herein may alternatively be obtained from a 3D modeled object representing the skin (i.e., outer surface) of a solid (e.g., corresponding to a B-Rep model representing the skin, i.e., exact surface), for example, by ray-casting on the 3D modeled object or tessellating the 3D modeled object. The tessellation may be performed in accordance with a rendering process of any 3D modeled object. Such a rendering process may be coded in any CAD system to display a graphical representation of the 3D modeled object. The 3D modeled object may be, or has been, designed by a user with a CAD system.

[0037] Providing the dataset may involve forming the dataset, for example, by creating 3D modeled objects. Alternatively, providing the dataset may involve retrieving the dataset from a (e.g., remote) memory where the dataset was created and stored. The 3D modeled objects in the dataset may all be 3D point clouds. Alternatively, they may all be 3D meshes. In such cases, the training method may include preprocessing to sample these meshes into 3D point clouds. The preprocessing may then involve centering each 3D mesh. The preprocessing may then involve independently rescaling the dimensions of each 3D mesh so that the vertices of the mesh fit exactly within a unit cube. The preprocessing may then involve extracting a dense point cloud from each normalized shape in the dataset, for example, by ray-casting the point clouds into six orthographic views. The preprocessing may then involve uniformly subsampling each point cloud (e.g., downsampling each point cloud to the same size). Subsampling may start from a random point in the point cloud and iteratively select the furthest point in the point cloud from an already selected point to arrive at the desired number of points.

[0038] The 3D modeled objects of the dataset may be plausible (e.g., realistic) 3D modeled objects. A plausible 3D modeled object may specify a 3D modeled object that represents a real-world object, e.g., a plausible mechanical part. A plausible mechanical part may specify a mechanical part that can be realistically manufactured in a real-world industrial manufacturing process. A plausible mechanical part may refer to a mechanical part that respects all constraints that must be respected to realistically manufacture the mechanical part in a real-world industrial manufacturing process. The constraints may consist of one or more of mechanical constraints (e.g., constraints arising from the laws of classical mechanics), functional constraints (e.g., constraints related to one or more mechanical functions performed by the mechanical part once manufactured), manufacturing constraints (e.g., constraints related to the ability to apply one or more manufacturing tools to the mechanical part during one or more manufacturing steps to produce the mechanical part), structural constraints (e.g., constraints related to the strength and / or resistance of the mechanical part), and / or assembly constraints (e.g., constraints defining how the mechanical part can be assembled with one or more other mechanical parts).

[0039] As previously mentioned, the 3D modeled objects of the provided dataset may all or substantially all belong to a single category of 3D modeled objects, or may form different categories of 3D modeled objects. In either case, the 3D modeled objects of the provided dataset represent respective objects from the real world, such as machine parts as previously discussed.

[0040] Explain neural network learning.

[0041] As is known in the field of machine learning, processing an input by a neural network involves applying an operation to the input, where the operation is defined by data including weight values. Therefore, training a neural network involves determining weight values ​​based on a dataset configured for training; such a dataset may be referred to as a learning dataset or training dataset. To this end, a dataset includes pieces of data, each of which forms a respective learning sample. The training samples represent a diversity of situations in which the neural network will be used after being trained. Any dataset referred to herein may include a number of training samples greater than 10,000, 100,000, or 1,000,000. Here, the neural network is trained on the provided dataset, meaning that the provided dataset is the learning / training dataset for the neural network. Training may be performed by any suitable known method.

[0042] The neural network is configured to infer a deformation basis for an input 3D modeled object. That is, the neural network receives a 3D modeled object as input and outputs a deformation basis for the 3D modeled object. To do this, the input 3D modeled object must be of the same data type as the 3D modeled object in the dataset. For example, if the 3D modeled object in the dataset is a 3D mesh, the input 3D modeled object is also a 3D mesh. Alternatively, if the 3D modeled object in the dataset is a 3D point cloud, the input 3D modeled object is also a 3D point cloud. Nevertheless, the 3D modeled object may be a 3D point cloud obtained by sampling a mesh. The deformation basis is a set of vectors, each of which is a deformation direction. The vectors may be linearly combined to deform the input 3D modeled object, and the linear combination, with its coefficients (also called amplitudes), results in the deformation. The vectors may be linearly combined in such a way that the deformed input 3D modeled object approaches the input 3D modeled object if the linear combination is sufficiently small. In such a case, the deformation is realistic. Mathematically, the deformation is realistic as long as the manifold is sufficiently close to the deformation on the tangent space defined by the deformation basis. The deformation basis is a basis in the sense of linear algebra, a set of linearly independent vectors, and the basis may be normalized, e.g., orthogonal. Specifically, the training aims for the neural network to infer, or at least tend to infer, a basis of deformation vectors that are linearly independent (e.g., uncorrelated and / or orthogonal vectors, as discussed further below). The basis of deformation vectors may have a fixed size (e.g., between 2 and 10 vectors, or greater than 10 vectors). In other words, the neural network may always (i.e., for each input 3D modeled object) output a deformation basis that has the same fixed number of vectors.If the dataset consists of plausible 3D modeled objects, the neural network will infer a deformation basis that has plausible deformation vectors of the 3D modeled objects, since it has been trained to do so.

[0043] The neural network has an architecture configured to take a 3D modeled object as input and output a deformation basis for the 3D modeled object. The neural network may be composed of an encoder and a deep feedforward neural network. The encoder is configured to receive a 3D modeled object as input and output a latent vector representing the input 3D modeled object. Thus, the encoder may be configured to receive a 3D mesh or a 3D point cloud (e.g., sampled from a 3D mesh) as input. The deep feedforward neural network is configured to receive the latent vector output by the encoder as input and output a deformation basis for the 3D modeled object represented by the latent vector. The encoder architecture may be based on PointNet (e.g., as described in the previously cited reference [8], which is incorporated herein by reference) or any extension thereof. Alternatively, the encoder architecture may use a mesh topology (i.e., when the 3D modeled objects in the dataset are 3D meshes), as done, for example, in mesh CNNs (see, e.g., the previously cited reference [9]).

[0044] The training may minimize a loss for at least a portion of the dataset, where for each 3D modeled object in the at least a portion of the dataset and for each candidate deformation basis having vectors, the loss penalizes a distance between a deformation of the 3D modeled object with a linear combination of the vectors and another 3D modeled object.

[0045] Discuss loss minimization.

[0046] At least a portion of the dataset is a sample of the 3D modeled object of the dataset. For example, at least a portion of the dataset may be a mini-batch. In that case, the training is performed mini-batch by mini-batch, and consists of minimizing a loss for each mini-batch. The concept of mini-batch training is known in the field of machine learning. For example, the training may implement a known mini-batch stochastic optimization method such as mini-batch stochastic gradient descent. Mini-batch training improves the efficiency of training, as is known per se from the field of machine learning.

[0047] For each 3D modeled object in at least some of the datasets, the training evaluates the deformation basis of the modeled object output by the neural network with the current weight values. This calculated deformation basis forms a candidate deformation basis for the 3D modeled object. The loss penalizes the distance between the vector of this candidate deformation basis and the deformation of the 3D modeled object resulting from a linear combination with another 3D modeled object. That is, the loss tends to be large when the distance between the deformation of the 3D modeled object resulting from the linear combination of the vectors of this candidate basis and another 3D modeled object is large. In this case, the training to minimize the loss modifies the weights of the neural network to reduce the loss value and obtain a new candidate. The training continues to do so (e.g., for each mini-batch) until the loss reaches its minimum value, or at least a sufficiently small value (e.g., with respect to a convergence criterion).

[0048] The other 3D modeled object may be a target object (i.e., a linear combination of vectors for deforming the 3D modeled object). This allows a neural network to be trained to infer the basis of the deformation. For example, the other 3D modeled object may be a 3D modeled object, among at least some of the 3D modeled objects in the dataset, that is close to at least some of the 3D modeled objects in the dataset. In other words, the learning may include selecting another 3D modeled object from among the at least some of the 3D modeled objects in the dataset, based on a distance from the at least some of the 3D modeled objects in the dataset. The selection may include calculating a distance between the at least some of the 3D modeled objects in the dataset and the at least some of the 3D modeled objects in the dataset, and evaluating which 3D modeled objects are closer, i.e., for which 3D modeled objects the distance is small. The selection may then select the 3D modeled object with the smallest distance. In an example, the selection may evaluate which 3D modeled object is closest, i.e., for which 3D modeled object the distance is smallest. In other words, the selection involves calculating a minimum distance and evaluating which 3D modeled object results in this minimum. In these examples, the other 3D modeled object is a 3D modeled object that is closest to the at least some 3D modeled object in the dataset among the 3D modeled objects in the dataset. In the examples, the learning is performed on a mini-batch basis, and for each mini-batch, the learning comprises minimizing a loss and selecting the closest 3D modeled object from among the 3D modeled objects in the mini-batch. In other words, the selection of the other 3D modeled object is performed from among the objects in the mini-batch, where the at least some of the dataset is the mini-batch.

[0049] The distances from at least some of the 3D modeled objects of the dataset may be any distance, such as distances between 3D modeled objects, e.g., distances in 3D between point clouds or 3D meshes, etc. In an example, the distances may be distances in the latent space, as currently discussed.

[0050] Specifically, in these examples, the learning method may comprise providing a separate encoder configured to take 3D modeled objects of a provided dataset as input and output latent vectors encoding the 3D modeled objects, e.g., prior to learning, e.g., in an early stage of the learning method. The separate encoder may be trained using an autoencoder framework, a classification task, or other machine learning method to learn a meaningful latent space of the 3D modeled objects, as known in the field of machine learning. Providing the separate encoder may comprise training the separate encoder. Alternatively, the separate encoder may be pre-trained. In these examples, the distance between the first 3D modeled object and the second 3D modeled object may be the distance in the separate encoder's latent space between the results of applying the separate encoder to the first 3D modeled object and the results of applying the separate encoder to the second 3D modeled object. This latent space distance allows a neural network to more accurately infer the basis of deformations, since the latent vectors output by the separate encoder capture the semantics of the encoded 3D modeled objects. In other words, the encoder implicitly clusters the dataset with respect to semantics. This allows another 3D modeled object to be determined as the 3D modeled object that is most similar to at least some of the 3D modeled objects in the dataset. The use of this latent space distance in this way can not only improve the accuracy of learning, but also improve efficiency compared to, for example, distances in the 3D modeled object space, which are computationally more expensive.

[0051] The selection of the alternative 3D modeled objects may occur prior to minimization, for example, the results of applying the alternative encoders to the 3D modeled objects of the dataset may be pre-computed, i.e., calculated prior to minimization.

[0052] A linear combination has coefficients, each corresponding to the contribution of each vector in the combination. The coefficients are calculated by the neural network during training, or they can be the result of optimizations performed during training.

[0053] The loss may penalize the minimum value of the distance between a deformation of a 3D modeled object by a linear combination of vectors and another 3D modeled object. Thus, when minimizing the loss, the optimal deformation by the linear combination, i.e., the deformation that minimizes the distance, is selected. The minimum value may be, for example, the minimum value of the distance between a deformation of a 3D modeled object by a linear combination of vectors and another 3D modeled object, or the minimum value of all possible values ​​of the coefficients of the linear combination. The loss penalizes the minimum value in that a larger minimum value tends to increase the loss. In this case, training to minimize the loss modifies the weights of the neural network to lower the loss value. Training continues (e.g., for each mini-batch) until the loss reaches a minimum, or at least a sufficiently small value (e.g., with respect to a convergence criterion).

[0054] The losses may be of the following types:

number

number

[0055] e1, ,e N may already be a 3D point cloud, in which case e1=p1,…,e N =p N Alternatively, e1, ,e N may be a 3D mesh, in which case p1,…,p N is e1,…,e N This may be due to sampling. CH d can be any 3D distance between 3D point clouds, such as earthmover distance or chamfer distance.CH Examples include the following types:

number

[0056] q1 is calculated as follows: Let f' be another encoder as described above. Then, for each i, its training method is

number

[0057] The loss may further penalize sparseness by including, for each 3D modeled object and each candidate deformation basis, a (e.g., minimum) function that takes as input the coefficients of the linear combination. The sparseness function may, for example, be the L1 norm of the coefficients of the linear combination, as previously discussed (see, e.g., Reference [6], incorporated herein by reference). In other words, when the sparseness-inducing function (e.g., minimum) is large, the loss tends to be large. In this case, training to minimize the loss modifies the weights of the neural network to lower the loss value. Training does so (e.g., for each mini-batch) until the loss reaches a minimum, or at least a sufficiently small value (e.g., with respect to a convergence criterion). In this case, the loss may be of the following type:

number

[0058] where λ is a trade-off parameter (a well-known concept in machine learning), φ is a sparsity-inducing function such as the L1 norm (see, e.g., previously cited reference [6], which is incorporated herein by reference), and (α1, ..., α p) is a vector of coefficients, also called the vector of amplitudes.

[0059] This allows the coefficients / amplitudes to be sparse during training. Ultimately, this allows the neural network to be trained to output uncorrelated deformation basis vectors, which further improves the generalization power of the training method. It is noteworthy that a neural network trained in this way can output a satisfactory basis because the sparseness allows it to use a small number of vectors to calculate the deformation of the input 3D modeled object as a linear combination of basis vectors. Furthermore, the output basis is, or at least tends to be, a basis in the linear algebraic sense.

[0060] The loss may further include, for each 3D modeled object and each candidate deformation basis, rewarding orthogonality of the candidate deformation bases. In other words, the loss may further include a term that captures the orthogonality of the candidate deformation bases of the 3D modeled object estimated by the neural network with the current weight values, and the term tends to be small if the deformation bases tend to be orthogonal. The loss in this case may be of the following type:

number

[0061] δ is a trade-off parameter (a well-known concept in machine learning) that further allows the neural network to learn to output transformed basis vectors that tend to be orthogonal matrices and therefore uncorrelated, such that such bases are, or at least tend to be, basis in the linear algebraic sense.

[0062] The loss may further comprise any suitable regularization function as known per se from the field of machine learning.

[0063] The implementation of the learning method is discussed.

[0064] The present embodiments bring improvements to the fields of 3D modeling and 3D machine learning. The results of the present embodiments can be used in fields such as virtual and augmented reality (more generally, any kind of immersive experience), video games, manufacturing and 3D printing, or 3D modeling. The present embodiments provide a solution to compute realistic sparse deformation linear bases for any 3D modeled object, which can be used to perform shape synthesis, shape reconstruction from images, or shape matching.

[0065] In this embodiment, we provide a plausible linear deformation basis for any modeled 3D object, such that any linear combination of the deformation vectors that make up the basis can realistically deform the model in a local neighborhood. Such deformations can be used in 3D modeling software, particularly 3D editing applications, for inferring large nonlinear deformations and shape matching.

[0066] In this embodiment, a neural network (train-while-training neural network) is trained to learn a linear 3D deformation field with sparse constraints to obtain plausible deformations that are independent of each other. In this embodiment, a neural network is trained (train-while-training neural network) to infer a deformation basis at each point of an input 3D modeled object from a provided dataset, a dataset of unlabeled 3D modeled objects. To train this network, training proceeds mini-batch by mini-batch, as described above.

[0067] (1) Match each 3D modeled object in each batch with the closest 3D modeled object in the same batch, i.e., select another 3D modeled object from the 3D modeled objects in the batch that is closest to the 3D modeled object in the batch. (2) Compute a linear combination of the predicted deformation vectors of the 3D modeled objects that minimizes the distance between the deformation vector inferred by the neural network with the current weights and the nearest 3D modeled object in the batch, with an additional sparsity regularization term. (3) Minimize the same loss (the distance between the deformed 3D modeled object and its nearest 3D modeled object) across the neural network weights, optimizing the basis for the predicted deformation.

[0068] In this embodiment, it is not necessary to cluster different types of 3D modeled objects and train them individually. Instead, it is suitable for direct training of the entire dataset. In this embodiment, uncorrelated deformation-based learning is performed, where each deformation is plausible in its own right. The estimated deformations are linear, making it possible to use them in real time. The deformations are estimated as a 3D deformation field, maximizing the expressiveness of the input 3D modeled object.

[0069] When providing a dataset, in this embodiment, if the provided dataset consists of 3D meshes, preprocessing may be performed as described above. Preprocessing consists of centering each mesh and independently rescaling the dimensions of each mesh so that the mesh vertices fit exactly within a unit cube. Next, preprocessing begins by extracting a dense point cloud from each normalized shape in the dataset by ray-casting it into six orthographic views. Next, preprocessing uniformly subsamples each point cloud (e.g., downsamples each point cloud to the same size). To do this, preprocessing may start with a random point in the point cloud and iteratively select the point in the point cloud that is farthest from the already selected point until a predefined (e.g., desired) number of points is reached.

[0070] The architecture of the neural network according to the currently discussed embodiment will now be discussed.

[0071] In this embodiment, the neural network infers the same fixed number of deformation basis vectors (e.g., for any input 3D modeled object). Let n be this fixed and same size of the inferred deformation basis. In this embodiment, the neural network receives a 3D point cloud or a 3D mesh as input and generates latent vectors

number

[0072] f w In addition to this, this implementation also uses a deep feedforward neural network g w Including the design of g w takes as input a 3D point x as well as a latent vector h of a 3D point cloud or a 3D mesh. w is the deformation basis at x

number

number

number

number

number

number

number

[0073] This implementation may include providing another encoder architecture f' similar to fw. As previously described, f' is trained in an autoencoder framework, or classification task, or any other machine learning method, to learn a meaningful latent space of 3D objects. In this embodiment, the other encoder is used to efficiently match 3D modeled objects (3D meshes or 3D point clouds) with their nearest 3D modeled objects in batches during training of fw.

[0074] In this embodiment, unsupervised learning is performed when training the neural network, which will now be discussed.

[0075] In training, both networks fw and gw are trained on a set of N points p1,...,p by any mini-batch stochastic optimization (e.g., mini-batch stochastic gradient descent). N It learns by minimizing the loss E(w) (also called "energy") for e1, ..., e N Let be the corresponding 3D modeled object in the input space of fw. That is, let these inputs be the point cloud e i =p i and if it is a mesh, e i HAp i is the source mesh used to sample.

[0076] For each i,

number

[0077] Then the loss is given by:

number

[0078] d CH is here the chamfer 3D distance, but can be replaced by other 3D losses such as the earthmover distance.

number

[0079] φ is L 1 is a sparsity-inducing function such as the norm (reference [6]). (α1,...,α p ) is called the vector of amplitudes. In this implementation, we seek to make the amplitudes sparse in order to make the deformations as uncorrelated as possible, so that any deformation vector is plausible on its own. For the same reason, this implementation enforces that the deformation basis is orthonormal with respect to the δ penalty, which precedes a term that captures the regularity of the basis as discussed previously.

[0080] σ(e i ,e i +v) may be any regularization function, including but not limited to combinations of the following: ·e i and e i +v gives a penalty for the difference in edge length between i If is a point cloud, this implementation involves generating edges on the point cloud using a k nearest neighbor graph). ·e iand e i +v (i.e., L(e i ) and L(e i +v), where L is e i is the Laplace operator of λ, δ, and γ are trade-off parameters.

[0081] To minimize E(w), training involves computing the gradient of E(w) for each input in the batch, as well as the amplitudes (α1,…, α p )

[0082] For each input i,

number

[0083] But the minimum

number

number

[0084] where:

number

number

number

[0085] The envelope theorem states that

number

number

number

number

number

[0086] FIG. 2 is a diagram illustrating a deep model architecture of a neural network trained according to the currently discussed embodiment.

[0087] Applications of learning methods will be discussed.

[0088] Application of a training method refers to a computer-implemented method of using a neural network that can be trained according to the training method, for example, a neural network that has been trained (i.e., trained) directly by the training method. The method of use may be configured to use the neural network to infer one or more deformation bases that each correspond to an input 3D modeled object, and to use the inferred one or more deformation bases. The training method may, for example, constitute the method of use as a further step of the training method that is performed after training.

[0089] A first example of the application of learning methods is depth frame reconstruction (see, for example, the previously cited reference [3], which is incorporated herein by reference). Depth frame reconstruction involves providing a 3D mesh and a depth map of another object. The other object is near the provided 3D mesh. Depth frame reconstruction then consists of applying a neural network to the 3D mesh to obtain a deformation basis for the 3D mesh. Depth frame reconstruction then uses this deformation basis to deform the provided 3D mesh into the other object represented by the depth map. Next, depth frame reconstruction further includes optimizing the deformation basis to fit the depth map; i.e., depth frame reconstruction consists of optimizing coefficients of a linear combination of deformation basis vectors so that deforming the 3D mesh by the linear combination fits the depth map. The goal of the reconstruction is to obtain a 3D modeled object corresponding to the depth map.

[0090] A second example of the application of the learning method is shape matching. Shape matching involves providing two similar 3D meshes e1 and e2. Then, shape matching calculates a deformation basis for the first 3D mesh e1 and optimizes the coefficients of a linear combination of the deformation basis vectors to deform the first 3D mesh e1 so that it matches the second 3D mesh e2. For example, shape matching may be calculated as follows:

number

number

[0091] Figure 3 shows a first example of shape matching: the first 3D mesh is a chair 30, the second 3D mesh is a chair 36, and chair 30 is to be matched. Chair 34 shows the result of the coefficient optimization, and chair 32 shows the intermediate chair calculated during the optimization.

[0092] Figure 4 shows a first example of shape matching: the first 3D mesh is chair 40, the second 3D mesh is chair 46, and chair 40 is to be matched. Chair 44 shows the result of the coefficient optimization, and chair 42 shows the intermediate chair calculated during the optimization.

[0093] A third exemplary application of the learning method is 3D editing, which may consist of providing a first 3D modeled object and, starting from the first 3D modeled object, iteratively transforming the first 3D modeled object into a second 3D modeled object by applying a neural network to the iterative transformation of the first 3D modeled object.

Claims

1. 1. A computer-implemented machine learning method comprising: providing a dataset of a 3D modeled object; training a neural network configured to infer a deformation basis of an input 3D modeled object, the deformation basis comprising a plurality of vectors forming a deformation of the input 3D modeled object, the plurality of vectors being linearly combined to obtain a deformation of the input 3D modeled object; and The neural network an encoder configured to receive as input a 3D modeled object and to output a latent vector representing the 3D modeled object; a deep feedforward neural network configured to receive as input the latent vectors output by the encoder and to output a deformation basis of the input 3D modeled object represented by the latent vectors; and The learning step includes, for at least a portion of the dataset, minimizing, for each candidate deformation basis having a plurality of vectors, for each 3D modeled object in the at least a portion of the dataset, a loss that penalizes a distance between a deformation of the 3D modeled object by a linear combination of the plurality of vectors and another 3D modeled object, the other 3D modeled object being a 3D modeled object that is closest to the 3D modeled object in the at least a portion of the dataset among the 3D modeled objects in the at least a portion of the dataset, and the other 3D modeled object being a target 3D modeled object using the linear combination of vectors to deform the 3D modeled object in the at least a portion of the dataset. Machine learning methods.

2. The training is performed mini-batch by mini-batch, and includes minimizing a loss for each mini-batch. The method of claim 1.

3. The training is performed for each mini-batch, and for each mini-batch, the training includes minimizing a loss and selecting the closest 3D modeled object from among the 3D modeled objects of the mini-batch.

3. The method according to claim 1 or 2.

4. The loss penalizes the minimum value of the distance between the vector and the deformation of the 3D modeled object by a linear combination of the vector and the other 3D modeled object.

4. The method according to any one of claims 1 to 3.

5. The loss is of the following type: [Equation 1] where: ・e 1 , ..., e N are 3D modeled objects in at least a portion of the dataset, and N is the number of objects in at least a portion of the dataset; ・p 1 , ..., p N is e 1 , ..., e N is the point cloud obtained from Each 3D modeled object e i For g w (f w (e i ),p i ) 1 , ..., g w (f w (e i ),p i ) n is a vector of candidate deformation bases for the 3D modeled object, n is the size of the candidate deformation bases, [Equation 2] is the deformation of the 3D modeled object ei by a linear combination of vectors, α1,...,αN are the coefficients of the linear combination, qi is a point cloud obtained from another 3D modeled object, ・d CH is the distance, The neural network has weights, where w represents the weights of the neural network, ・f w receives a 3D modeled object as input and an encoder configured to output a latent vector representing the object; ・g w takes the latent vector output by the encoder as input and is a deep feedforward neural network configured to output deformation bases for a 3D modeled object represented by a The method of claim 4.

6. The loss further penalizes the sparsity-inducing function that takes as input coefficients the linear combination.

6. The method according to any one of claims 1 to 5.

7. The loss further rewards orthogonality of the candidate transformation bases.

7. The method according to any one of claims 1 to 6.

8. A neural network trained according to the method of any one of claims 1 to 7, comprising: A neural network that causes a computer to accept as input a 3D modeled object and to output an inferred deformation basis for said 3D modeled object.

9. comprising instructions which, when executed on the computer, cause the computer to carry out a method according to any one of claims 1 to 7. Computer program.

10. A computer-readable data storage medium having recorded thereon the computer program of claim 9 and / or the neural network of claim 8.

11. A computer comprising a processor coupled to a memory, the memory storing the computer program of claim 9 and / or the neural network of claim 8.

Citation Information

Patent Citations

  • Selecting method for shoes

    JP2000090272A

  • Morphing image generation device, and morphing image generation method

    JP2019096130A

  • Prosthesis shape data generation system

    WO2019103010A1