Convolutional neural network structure for learnable 3D shape learning, and method and system for 3D shape learning using the same
The convolutional neural network model extracts geodesic and geometric features from 3D shape data using learnable descriptors and 1D convolution, addressing noise and complexity issues while providing explainable classification and visualization.
Patent Information
- Application Number
- JP2025502697
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-09
- Filing Date
- 2023-08-03
- Publication Date
- 2025-07-17
- Estimated Expiration
- 2043-08-03
AI Technical Summary
Existing 3D shape learning methods using deep learning are sensitive to noise and computational complexity due to data conversion, and lack explainability, especially when processing irregular shapes, and existing explainable AI techniques are not applicable to 3D shapes.
A convolutional neural network model that extracts geodesic and geometric features from 3D shape data using Order-Invariant Kernel Mapping and learnable descriptors, performing 1D convolution operations to maintain spatial information and enable explainable classification.
Improves object classification performance and enables visualization of inference results without modifying the model, maintaining spatial information and providing interpretability for 3D shape data.
Smart Images

Figure 2025523189000001_ABST
Abstract
Description
Technical Field
[0001] The following description relates to 3D shape learning technology.
Background Art
[0002] 3D shape analysis using deep learning mainly consists of multi-view and voxel methods. Recently, there has been increasing interest in models using graph and mesh methods.
[0003] In order to process multi-view and voxel-based models with neural networks, irregular 3D shapes are converted into regular expressions. Such methods can perform general convolution operations, but are sensitive to noise due to the loss of topological connectivity and the reduction of data density during the expression conversion, resulting in unstable results or requiring high computational complexity.
[0004] The graph method performs convolution operations in non-Euclidean spaces using geodesic-based features such as the path between nodes and edges of the graph. The mesh method performs convolution operations adapted to 3D shapes using geometric features such as the position and direction of nodes and faces. Such methods empirically use low-level features to process specific 3D shape datasets, so there are limitations in processing all types of 3D shapes.
[0005] Furthermore, deep learning cannot explain the reasons for inferences due to its internal non-linearity. To solve this problem, explainable artificial intelligence (XAI) has been actively studied. However, although many studies have been conducted on model explanation techniques for 2D images, general methods cannot be applied to 3D shapes, and only techniques specialized for models and data have been gradually studied.
Summary of the Invention
Problems to be Solved by the Invention
[0006] Provided are a method and a system for extracting geodesic features and geometric features from 3D shape data and classifying an object by a convolution operation on the extracted geodesic features and the extracted geometric features based on the surfaces constituting the 3D shape data.
[0007] Provided is a convolutional neural network model for explainable 3D shape learning that perceives geodesic features and geometric features of 3D shape data while maintaining spatial information.
Means for Solving the Problems
[0008] A method for shape learning executed by a shape learning system may include a step of extracting geodesic features and geometric features from 3D shape data, and a step of classifying an object by a convolution operation on the extracted geodesic features and the extracted geometric features based on the surfaces constituting the 3D shape data.
[0009] A method for shape learning may further include a step of visualizing the result of classifying an object from 3D shape data using a convolutional neural network model for explainable 3D shape learning.
[0010] A convolutional neural network model for explainable 3D shape learning may be composed of a descriptor layer that extracts geodesic features and geometric features from 3D shape data using each target surface and adjacent surfaces, and a convolutional layer that performs a convolution operation using the extracted geodesic features and geometric features.
[0011] The 3D shape data may be composed of nodes, edges, and surfaces that constitute the 3D shape data having an irregular structure.
[0012] The extraction step may include a step of generating a data structure to obtain information on the nodes and edges of each surface based on the identification information configured in a list form for the surfaces adjacent to each surface in units of each surface for the surfaces constituting the three-dimensional shape data.
[0013] The extraction step may include a step of extracting geodesic features from the three-dimensional shape data using Order-Invariant Kernel Mapping.
[0014] The extraction step may include a step of extracting geometric features from the three-dimensional shape data based on the positions and directions of the target surface and the surrounding surfaces.
[0015] The extraction step may include a step of extracting geometric features from the three-dimensional shape data based on the distance ratio and the difference angle information between the target surface and the surrounding surfaces.
[0016] The classification step may include a step of inputting a feature vector for each surface integrating the geodesic features and the geometric features into a convolutional layer.
[0017] The classification step may include a step of performing 1D convolution using a feature vector inherited by each surface based on the feature vector for each surface and a feature vector aggregating the relationship between a specific surface and the adjacent surfaces of the specific surface.
[0018] The extraction step may include a step of performing a GAP (Global Average Pooling) operation by outputting a result so that the number of surfaces is maintained through the executed 1D convolution to obtain a score for each class, and classifying an object based on the obtained score.
[0019] A computer program may be included which is recorded on a non-transitory computer-readable recording medium so as to cause the above-described shape learning system to execute a method for shape learning.
[0020] The shape learning system may include a feature extraction unit that extracts geodesic features and geometric features from three-dimensional shape data, and an object classification unit that classifies an object by performing a convolution operation on the extracted geodesic features and the extracted geometric features based on the surfaces constituting the three-dimensional shape data.
[0021] The shape learning system may visualize the result of classifying an object from three-dimensional shape data by using a convolution-based neural network model for explainable three-dimensional shape learning.
Advantages of the Invention
[0022] By extracting geodesic features and geometric features from three-dimensional shape data and performing a convolution operation on the extracted geodesic features and the extracted geometric features based on the surfaces constituting the three-dimensional shape data, the object classification performance can be improved.
[0023] By using a convolution-based neural network model for explainable three-dimensional shape learning, the inference result can be visualized without modifying the model.
Brief Description of the Drawings
[0024]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
[0025] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings.
[0026] FIG. 1 is a block diagram for explaining the configuration of a shape learning system in one embodiment, and FIG. 2 is a flowchart for explaining a method for shape learning in one embodiment.
[0027] The processor of the shape learning system 100 may include a feature extraction unit 110 and an object classification unit 120. Components of such a processor may be representations of different functions that are executed by the processor according to control instructions provided by program code recorded in the shape learning system. The processor and the components of the processor may control the shape learning system to execute steps 210 to 220 included in the method for shape learning in FIG. 2. At this time, the processor and the components of the processor may be realized to execute instructions by the code of the operating system included in the memory and the code of at least one program.
[0028] The processor may load the program code recorded in the program file for the method for shape learning into the memory. For example, when the program is executed in the shape learning system, the processor may control the shape learning system to load the program code from the program file into the memory according to the control of the operating system. At this time, the feature extraction unit 110 and the object classification unit 120 may be different functional expressions of the processor for executing the instructions of the corresponding parts of the program code loaded into the memory to execute the subsequent steps 210 to 220.
[0029] In step 210, the feature extraction unit 110 may extract geodesic features and geometric features from the three-dimensional shape data. In graph theory, geodesic information represents the distance between two nodes by the shortest path connecting the two points. Since adjacent faces (hereinafter referred to as "adjacent faces") in the three-dimensional shape data have the same nodes and edges, they will have the same geodesic information. When the data is represented on a vector space, the nodes and edges may be represented by vectors and operations may be performed. There is a geodesic convolution as an operation method designed by utilizing the point of representing nodes and edges by vectors and performing operations, but unstable results may occur depending on the order of the nodes and edges. Accordingly, the feature extraction unit 110 may extract geodesic features from the three-dimensional shape data by using Order-Invariant Kernel Mapping. The feature extraction unit 110 can extract geodesic features without loss of spatial information by performing operations on the information of the nodes and edges of the target face and the surrounding faces regardless of the order. In addition, the feature extraction unit 110 may extract geometric features from the three-dimensional shape data based on the positions and directions of the target face and the surrounding faces. The feature extraction unit 110 can obtain a robust result independent of the data size by performing operations regardless of the order based on the distance ratio and the difference angle information between the target face and the surrounding faces.
[0030] In stage 220, the object classification unit 120 may classify the object by performing a convolution operation on the extracted geodesic features and the extracted geometric features based on the surfaces constituting the three-dimensional shape data. The object classification unit 120 may integrate the information of the target surface and the adjacent surfaces by surface unit operations. The object classification unit 120 may apply the model description technique used in the classification problem of two-dimensional images by using one-dimensional convolution operations.
[0031] FIG. 3 is a diagram for explaining the operation of an explainable model in one embodiment.
[0032] Three-dimensional shape data is data actively used in computer graphics. However, due to the characteristic that the data is irregular, it is difficult to learn with a neural network. The technology of converting three-dimensional shape data into a simple structure for processing is difficult to explain the inference of the neural network due to the loss of spatial information. In the embodiment, the purpose is to efficiently process the data through an explainable model (ExMeshCNN) 300 that senses the geodesic and geometric characteristics of three-dimensional shape data while maintaining spatial information, and visualize the explanation.
[0033] The explainable model 300 means a convolutional-based neural network model for explainable three-dimensional shape learning. The explainable model 300 may be composed of a descriptor layer 310 and a convolutional layer 320. The descriptor layer 310 may include a geodesic descriptor 311 that extracts geodesic features using each target surface and adjacent surfaces from the three-dimensional shape data, and a geometric descriptor 312 that extracts geometric features using each target surface and adjacent surfaces from the three-dimensional shape data. The convolutional layer 320 may perform a convolution operation using the extracted geodesic features and geometric features.
[0034] In the first layer of the explainable model 300, a learnable edge-based geodesic descriptor 311 and a face-based geometric descriptor 312 may be attached to obtain high-level features from the surface realized by 1D-CNN operations. Next, in the convolutional layer 320 that performs convolution on each surface, following the descriptor layer 310, the goal is to learn the regional features within adjacent surfaces. Finally, the GAP operation and the softmax layer are configured to make the output be logits. As a result, the explainable model 300 applies the input attribute method by learning the mesh in an end-to-end manner while maintaining spatial information from input to output, allowing for the interpretability of the model.
[0035] The shape learning system may utilize the explainable model 300 to classify objects from 3D shape data. At this time, since the 3D shape data has an irregular structure of nodes, edges, and faces, it cannot be used as the input of a neural network model. Therefore, the shape learning system may perform data preprocessing. Specifically, the shape learning system may generate, in list form, the identification information for three adjacent faces for each face of the 3D shape data. The shape learning system may generate a data structure based on the identification information so as to be able to obtain information regarding the nodes and edges of each face.
[0036] FIG. 4 is a diagram for explaining the geodesic path between a target face and adjacent faces in one embodiment.
[0037] The shape learning system can extract geodesic features and geometric features from 3D shape data and classify objects by convolution operations based on the faces that make up the 3D shape data. In addition, the shape learning system is designed so that general techniques for explaining 2D image classification models can be easily applied in terms of dimensions, enabling the explanation of the behavior of the model.
[0038] More specifically, the descriptor layer is the first layer that attempts to mine geodesic and geometric features from the faces that make up the mesh. Here, two types of learnable descriptors are proposed to capture geodesic and geometric features.
[0039] The geodesic descriptor aims to extract the geodesic regional features of a face, which have sometimes been ignored because it is difficult for conventional mesh-based models to define the geodesic relationship between two faces. Generally, the geodesic distance between two vertices is defined by the number of edges on the shortest path connecting the two vertices. Since the shape learning system performs feature extraction and convolution at the face level, the relationship between faces can be shown by using the geodesic distance. However, in order to simply solve the problem when applying the geodesic distance between faces, a geodesic descriptor is introduced.
[0040] Let the target face be JPEG2025523189000002.jpg87. Next, To estimate the geodesic relationship with the face adjacent to JPEG2025523189000003.jpg87, Consider the vertices and edges aligned with JPEG2025523189000004.jpg87. Let the center point of JPEG2025523189000005.jpg87 be JPEG2025523189000006.jpg78, and Let any one of the three vertices of JPEG2025523189000007.jpg87 be j ∈ {1, 2, 3} JPEG2025523189000008.jpg99. Then, The relationship between JPEG2025523189000009.jpg78 and JPEG2025523189000010.jpg99 is JPEG2025523189000011.jpg924. Also, JPEG2025523189000012.jpg99 (k ∈ 1, 2) is Excluding the vertices belonging to JPEG2025523189000013.jpg87, Among the adjacent faces of JPEG2025523189000014.jpg87 Let the vertices be geodesically connected with JPEG2025523189000015.jpg99. JPEG2025523189000016.jpg99 and Meaning the relationship with JPEG2025523189000017.jpg99 JPEG2025523189000018.jpg1129 is shown. Eventually, the target face JPEG2025523189000019.jpg87 and its adjacent face (that is, JPEG2025523189000020.jpg98) the geodesic distance between The center point of JPEG2025523189000021.jpg87 JPEG2025523189000022.jpg78 and It becomes the number of edges on the shortest path between the vertices belonging to JPEG2025523189000023.jpg98.
[0041] However, one problem occurs. As shown in Figure 4, For JPEG2025523189000024.jpg87 and the adjacent face, the length of the shortest path becomes 2. For example, the center point JPEG2025523189000025.jpg78 and JPEG2025523189000026.jpg910, and the center point JPEG2025523189000027.jpg78 and The length of the shortest geodesic path between JPEG2025523189000028.jpg910 is 2, and it is the same in fact in all other cases. Eventually, the geodesic distance will not contain discriminative and significant information as a surface feature. As an alternative to the geodesic distance, geodesic convolution that extracts features while additionally considering the curvature of the path can be considered. This can solve the problem as described above because even if the lengths are the same, it differs depending on the curvature of the shortest path (for example, path 1 and path 2 shown in Figure 1). In geodesic convolution, the vertex The normal vector of JPEG2025523189000029.jpg99 is the center point from JPEG2025523189000030.jpg78 to the middle of the path to JPEG2025523189000031.jpg99 and is necessary to consider the curvature. The normal vector JPEG2025523189000032.jpg917 can be easily obtained by averaging the normal vectors of the surfaces surrounding the vertex JPEG2025523189000033.jpg99.
[0042] Unfortunately, there is some ambiguity in the geodesic convolution of the geodesic path due to the property that the order is not defined. For example, as shown in Figure 4, JPEG2025523189000034.jpg87 has two adjacent vertices JPEG2025523189000035.jpg910 and JPEG2025523189000036.jpg98, and generates paths through ( JPEG2025523189000037.jpg1015) and ( JPEG2025523189000038.jpg1015). At this time, since the order of the two paths is not defined, the result of the convolution operation including the two paths will be different each time.
[0043] To solve such a problem, instead of the geodesic distance and geodesic convolution, Propose the following descriptors that map JPEG2025523189000039.jpg65 to the feature space.
[0044]
Number
[0045] Figure 5 is a diagram for explaining the geometric relationship between the target surface and the adjacent surfaces in one embodiment.
[0046] Geometric features widely used for the surfaces of mesh data include the surface normal vector, the center, and the edge angle. However, since these are somewhat heuristic and low - level features, propose learnable descriptors that can automatically capture high - level representations.
[0047] The center point indicating the position of JPEG2025523189000046.jpg87 is JPEG2025523189000047.jpg78, and the center points of its adjacent surfaces are Set it to JPEG2025523189000048.jpg912. In this case, Set JPEG2025523189000049.jpg to 78 The Euclidean distance between JPEG2025523189000050.jpg912 may be a low-level feature. Next, The normal vector of JPEG2025523189000051.jpg87 The normal vector of the adjacent surface of JPEG2025523189000052.jpg88 Use the cross product to obtain the angle between two surfaces from the normal vectors of JPEG2025523189000053.jpg99 and JPEG2025523189000054.jpg88× JPEG2025523189000055.jpg99 may be calculated, which may also be a low-level feature. In addition to low-level features, high-level features may be obtained using the following kernel mapping.
[0048]
Number
[0049] Therefore, the following descriptor that is not affected by the order is proposed.
[0050]
Number
[0051] In summary, the first layer of the explainable model (ExMeshCNN) is for each target surface JPEG2025523189000061.jpg87 and its adjacent surface, and is composed of a geodesic descriptor and a geometric descriptor ( JPEG2025523189000062.jpg1115 and JPEG2025523189000063.jpg814) realized by 1D convolution operation. Next, the results of the two descriptors can be concatenated and input to the next convolutional layer.
[0052]
Number
[0053] Figure 6 is a diagram for explaining the structure of the convolutional layer of the explainable model in one embodiment.
[0054] Let N be the number of faces constituting the input mesh. After passing through the descriptor layer, the mesh may be represented by a matrix composed of face feature vectors JPEG2025523189000065.jpg928. Define JPEG2025523189000066.jpg939 as the dimension of the face feature vector (or the number of input channels of the next CNN). By doing so, the size of the matrix becomes N×C1.
[0055] Next, a series of 1D-CNN layers through which the output of the descriptor layer passes may be designed for the interpretable model. For each layer, the input matrix may be extended by a combination of two types of feature vectors JPEG2025523189000067.jpg65 and JPEG2025523189000068.jpg85 may be extended by a combination. Specifically, JPEG2025523189000069.jpg811 and JPEG2025523189000070.jpg78 may be defined as follows.
[0056]
Number
[0057] Next, 1-D convolution specialized for the above input format may be performed. Let JPEG2025523189000076.jpg942 be an arbitrary convolution filter. Here, the activation of each filter JPEG2025523189000077.jpg914 may be calculated as follows.
[0058]
Number
[0059] In such a way, each l-th layer consumes the input (N×C l ), expands this input to the size (2N×C l ), and then converts it to the output size of (2N×C l+1 ) by a 1-D convolution filter. As a result, the number of channels changes for each layer, but the width is maintained in the same way as N which is the number of sides. By this mechanism, the prominence of each side for decision-making can be easily understood.
[0060] In the last CNN layer, the number of output channels must be the same as the number of classes K. Next, scores for each class may be obtained by a Global Average Pooling (GAP) operation, and finally, classification may be performed based on the scores. GAP not only helps significantly reduce the number of parameters while maintaining accuracy, but also makes it easier to back-project the contribution of each aspect to the classification.
[0061] Figures 7 to 10 are diagrams for explaining examples of the results of an explainable model in one embodiment.
[0062] The structure of the explainable model is regarded as convolutional at the surface level and a final GAP operation that preserves the number of input surfaces. Despite not being specialized for 3D mesh data, it can emphasize a remarkable representation for classification by collaborating with conventional visual attribute methods. In Figures 7 to 10, we will show how representative visual attribute methods such as Layer-wise Relevance Propagation (LRP) and Gradient-weighted Class Activation Mapping (Grad-CAM) operate in the explainable model.
[0063] First, LRP will be explained. LRP adopts an explanation by a decomposition strategy. After regarding the final output of the model as a relevance score, the contribution of each neuron to the relevance score is back-propagated from the output layer to the input layer. At this time, there is a constraint that the sum of the contribution degrees of the neurons belonging to each layer must be the same for all layers. To apply LRP, the relevance JPEG2025523189000086.jpg810 may be assigned to the softmax output as follows.
[0064]
Number
[0065]
Number
[0066]
Number
[0067] Finally, in the first descriptor layer, the relevance may be distributed to the relevance scores at the plane level shown by JPEG2025523189000097.jpg97 respectively as follows.
[0068]
Number
[0069] Next, Grad-CAM is described. While LRP attempts to redistribute the relevance scores of the output into the input space, Grad-CAM focuses on the information in the last layer. Specifically, after calculating the gradients for each class flowing into the last convolutional layer, it generates a rough localization map indicating the important regions of the input.
[0070] A k ∈R N is the feature map of the k-th channel in the last CNN layer. By doing so, the weight value of the neuron importance for each class c th can be calculated as follows. JPEG2025523189000101.jpg810
[0071] JPEG2025523189000102.jpg1538 This captures the importance of the feature map of the k-th channel for the target class c. JPEG2025523189000103.jpg1115 means the GAP operation. Note that the size of the feature map N is the same as the number of faces.
[0072] Finally, the class-discriminative localization map of size N for the target class c represented by JPEG2025523189000104.jpg108 can be calculated as follows.
Equation
[0073] According to the embodiment, ModelNet40, Manifold40, SHREC11, and Cube are used as datasets for evaluating the classification accuracy of the explainable model, and COSEG is used as a dataset for evaluating the segmentation performance. Manifold40 is a dataset with the 2-Manifold condition added to ModelNet40, and SHREC11 is a small dataset with the training data set to 10 and 16 respectively for testing. Cube is a dataset in the form of an object thinly cut on one face of a hexahedron.
[0074] Tables 1 to 4 show the performance of the models for each data. Tables 1 and 2 compare the classification performance with conventional models, Table 3 compares the segmentation performance with conventional models, and Table 4 compares the number of parameters with conventional models.
[0075] To compare the performance of the method proposed in the embodiment with conventional methods, it was compared with graph and mesh-based models using similar methods. As can be seen from the experimental results, the method proposed in this embodiment showed high accuracy in a number of datasets compared to other conventional methods. It also showed high accuracy in segmentation. It was confirmed that the model proposed in the embodiment showed excellent performance despite being the lightest model among the models shown in the following table.
[0076]
Table 1
[0077]
Table 2
[0078]
Table 3
[0079]
Table 4
[0080] Figures 7 to 10 show the results of applying LRP and Grad-CAM, which are generally used in the description of two-dimensional image classification models, to an explainable model. Conventional methods deformed the model as a whole or used additional methodologies to apply the explanation technique. However, in the embodiments, the inference results of the model can be explained using the explanation technique without any deformation. The red color shown in Figures 7 to 10 is an important part of the inference result of the model.
[0081] Figure 7 shows the results obtained by executing LRP with a trained explainable model, and Figure 8 shows the results of applying Grad-CAM to the explainable model. Also, Figures 9 and 10 show the results of rotating the same mesh at various angles to observe all the surface protrusions displayed on the three-dimensional mesh. It can be reconfirmed that the explainable model learned some discriminatory features to make a final decision, and LRP and Grad-CAM successfully captured the features. For example, when observing a car from below, it can be seen that the important feature is the tire, and when looking at it from the front, it can be seen that the windshield is also an important feature. In the case of a bed, the frame and scattered quilts can be cited as important features, but the flat floor of the bed does not seem to be considered. As a result, it can be confirmed that the explainable model appropriately captures class-discriminatory features from the mesh data, and visualization can be easily realized by executing LRP and Grad-CAM with the explainable model.
[0082] The apparatuses described above may be implemented by hardware components, software components, and / or combinations of hardware components and software components. For example, the apparatuses and components described in the embodiments may be implemented using one or more general-purpose computers or special-purpose computers, such as a processor, a controller, an ALU (arithmetic logic unit), a digital signal processor, a microcomputer, an FPGA (field programmable gate array), a PLU (programmable logic unit), a microprocessor, or various apparatuses capable of executing instructions and responding. The processing device may execute an operating system (OS) and one or more software applications running on the OS. Further, the processing device may access data, record, manipulate, process, and generate data in response to the execution of the software. For the sake of convenience of understanding, it may be described as if one processing device is used, but those skilled in the art will understand that the processing device may include a plurality of processing elements and / or a plurality of types of processing elements. For example, the processing device may include a plurality of processors or one processor and one controller. Also, other processing configurations such as parallel processors are possible.
[0083] The software may include a computer program, code, instructions, or a combination of one or more of these, and may configure the processing device to operate as desired or command the processing device independently or collectively. The software and / or data may be embodied in any type of machine, component, physical device, virtual device (virtual equipmet), computer recording medium, or device for interpreting by the processing device or providing instructions or data to the processing device. The software may be distributed over a computer system connected by a network and recorded or executed in a distributed state. The software and data may be recorded on one or more computer-readable recording media.
[0084] The method according to the embodiment may be realized in the form of program instructions executable by various computer means and recorded on a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures, etc. alone or in combination. The program instructions recorded on the medium may be those specially designed for the embodiment or those that can be used and are known to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy (registered trademark) disks, and magnetic tapes, optical media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to record and execute program instructions such as ROMs, RAMs, and flash memories. Examples of program instructions include not only machine language codes such as those generated by compilers but also high-level language codes executable by a computer using an interpreter or the like.
[0085] As described above, the embodiment has been described based on limited embodiments and drawings. However, those skilled in the art will be able to make various modifications and variations from the above description. For example, even if the described technology is executed in a different order from the described method, and / or the components of the described system, structure, device, circuit, etc. are combined or assembled in a different form from the described method, or opposed or replaced by other components or equivalents, appropriate results can be achieved.
[0086] Therefore, even if they are different embodiments, as long as they are equivalent to the claims, they belong to the scope of the appended claims.
Claims
1. A method for shape learning executed by a shape learning system, comprising: extracting geodesic features and geometric features from three-dimensional shape data; classifying an object by a convolution operation on the extracted geodesic features and the extracted geometric features based on the surfaces constituting the three-dimensional shape data. A method for shape learning including the above steps.
2. Visualizing the result of classifying an object from the three-dimensional shape data by using a convolutional neural network model for explainable three-dimensional shape learning. The method for shape learning according to Claim 1, further including the above step.
3. The convolutional neural network model for explainable three-dimensional shape learning includes: a descriptor layer that extracts geodesic features and geometric features from the three-dimensional shape data by using each target surface and adjacent surfaces; and a convolutional layer that performs a convolution operation by using the extracted geodesic features and geometric features. The method for shape learning according to Claim 2.
4. The three-dimensional shape data is characterized in that the nodes, edges, and surfaces constituting the three-dimensional shape data have an irregular structure. The method for shape learning according to Claim 1.
5. The extracting step includes: generating a data structure to obtain information about the nodes and edges of each surface based on the identification information of a plurality of adjacent surfaces in a list form for each surface of the surfaces constituting the three-dimensional shape data. The method for shape learning according to Claim 1, including the above step.
6. The extracting step includes: extracting geodesic features from three-dimensional shape data by using order-invariant kernel mapping. The method for shape learning according to Claim 1, including the above step.
7. The extracting step includes: extracting geometric features from the three-dimensional shape data based on the positions and directions of the target surface and surrounding surfaces. The method for shape learning according to Claim 1, including the above step.
8. The extracting step includes: extracting geometric features from the three-dimensional shape data based on the distance ratio and difference angle information between the target surface and surrounding surfaces. The method for shape learning according to Claim 7, including the above step.
9. The classifying step includes: The step of inputting the feature vector for each surface integrating the geodesic feature and the geometric feature into a convolutional layer The method for shape learning according to claim 1, including this step
10. The step of classifying includes The step of performing 1D convolution using a feature vector inheriting each surface based on the feature vector for each surface, and a feature vector aggregating the relationship between the specific surface and the adjacent surfaces of the specific surface The method for shape learning according to claim 9, including this step
11. The step of extracting includes By outputting the result so that the number of surfaces is maintained by the performed 1D convolution, performing a GAP (Global Average Pooling) operation to obtain a score for each class, and performing object classification based on the obtained score The method for shape learning according to claim 10, including this step
12. A computer program recorded on a non-transitory computer-readable recording medium to cause the shape learning system to execute the method for shape learning according to any one of claims 1 to 11
13. A shape learning system, comprising A feature extraction unit that extracts geodesic features and geometric features from three-dimensional shape data, and An object classification unit that classifies an object by a convolution operation on the extracted geodesic features and the extracted geometric features based on the surfaces constituting the three-dimensional shape data The shape learning system, including this
14. The shape learning system Is characterized by visualizing the result of classifying an object from the three-dimensional shape data using a convolutional neural network model for explainable three-dimensional shape learning The shape learning system according to claim 13
Citation Information
Patent Citations
Three-dimensional model feature extraction method using patch convolution
CN113570692A
3D face recognition
JP2006502478A