Three-dimensional model retrieval method and system based on Laplacian features and attention
Through the Laplace-Beltrami feature function and attention mechanism, multiple sets of two-dimensional projected views are generated and view features are fusion, which solves the problems of loss of structural information and insufficient relationship between views in three-dimensional model retrieval, and improves the accuracy and stability of the retrieval.
Patent Information
- Application Number
- CN202510557581.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-18
AI Technical Summary
There are problems in the existing three-dimensional model retrieval methods such as loss of three-dimensional structure information and insufficient modeling of relationships between views, resulting in insufficient retrieval accuracy and stability.
A three-dimensional model retrieval method based on Laplace-Beltramie feature function and attention mechanism is adopted. By generating multiple sets of two-dimensional projected views and adaptively fusion view features based on attention mechanism, combined with view grouping strategy, the final three-dimensional model descriptor is formed.
It effectively alleviates the problem of three-dimensional structural information loss, enhances the robustness and discriminantity of model representation, and improves the accuracy and stability of search results.
Smart Images

Figure CN120336570A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of 3D model retrieval. Specifically, it relates to a 3D model retrieval method and system based on Laplace features and attention. Background Technique
[0002] The statements in this part only provide background technical information related to the present disclosure and do not necessarily constitute prior art.
[0003] 3D model retrieval has wide applications in fields such as industrial design, autonomous driving, and biomedicine. For a given query model, 3D model retrieval aims to quickly and accurately retrieve similar models from a large-scale 3D model library. How to quickly and accurately retrieve the target model similar to the query model in a huge model database has become an important basic ability to improve the efficiency and intelligence level of intelligent systems. Therefore, studying efficient and accurate 3D model retrieval methods has important practical significance and research value.
[0004] Existing 3D model retrieval methods can be roughly divided into four categories according to the model representation form: point cloud-based, voxel-based, mesh-based, and view-based methods. Currently, view-based methods project a 3D model from multiple perspectives to obtain a series of 2D projection views, and use deep learning to extract view features to represent the 3D model. Thanks to the rapid development of the image processing framework, such methods perform excellently in terms of retrieval accuracy and efficiency. However, there are still several key problems in view-based 3D model retrieval methods: (1) The problem of loss of 3D structure information: In the process of rendering a 3D model into 2D views, it is difficult to completely retain the geometric shape and spatial structure information of the model, and 2D images are difficult to fully express the spatial structure information of the 3D model, resulting in the compression or even loss of the model's spatial relationship and geometric details in the projection, affecting the overall retrieval accuracy; (2) The problem of insufficient modeling of the relationship between views: Existing methods generally adopt the method of independently extracting the features of each view and fail to fully capture the context correlation and complementary information between views, restricting the overall semantic expression ability of the overall model descriptor. (3) Coarse view fusion strategy: In the feature fusion stage, simple average pooling or splicing operations are generally used, which is difficult to model the importance difference of different views in space, resulting in problems such as view redundancy and unbalanced feature expression. Summary of the Invention
[0005] To solve the problems of easy loss of three-dimensional structure information and insufficient modeling of the relationship between views in the above-mentioned three-dimensional model retrieval, the present disclosure proposes a three-dimensional model retrieval method and system based on Laplace features and attention. Based on the feature fusion and association mechanism between views, the potential relationship between views is mined to enhance the robustness and discriminability of the model representation, and on the basis of ensuring the complete expression of the geometric features of the non-rigid model, the accuracy and stability of the retrieval results are improved.
[0006] To achieve the above object, the present disclosure adopts the following technical solutions: One or more embodiments provide a three-dimensional model retrieval method based on Laplace features and attention, including the following steps: Use the Laplace-Beltrami eigenfunction to represent the three-dimensional model to be retrieved; Render the low-frequency Laplace-Beltrami eigenfunction obtained by expressing the three-dimensional model from multiple perspectives to generate multiple sets of two-dimensional projection views; Extract features from the two-dimensional projection views, and adaptively fuse the feature information of the current view and adjacent views based on the attention mechanism to obtain the multi-perspective features after extraction; Group the multi-perspective features, fuse the view features within the group to form group-level features, aggregate the group-level features to form the model descriptor of the final three-dimensional model, and match the closest three-dimensional model based on the model descriptor.
[0007] One or more embodiments provide a three-dimensional model retrieval system based on Laplace features and attention, including: A Laplace-Beltrami feature calculation module configured to use the Laplace-Beltrami eigenfunction to represent the three-dimensional model to be retrieved; A feature function projection view generation module configured to render the low-frequency Laplace-Beltrami eigenfunction obtained by expressing the three-dimensional model from multiple perspectives to generate multiple sets of two-dimensional projection views; Adjacent view feature aggregation based on attention, configured to extract features from the two-dimensional projection views, and adaptively fuse the feature information of the current view and adjacent views based on the attention mechanism to obtain the multi-perspective features after extraction; A view feature grouping and aggregation module configured to group the multi-perspective features, fuse the view features within the group to form group-level features, aggregate the group-level features to form the model descriptor of the final three-dimensional model, and match the closest three-dimensional model based on the model descriptor.
[0008] An electronic device includes a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the steps in the above-mentioned three-dimensional model retrieval method based on Laplace features and attention are completed.
[0009] A computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, the steps in the above-mentioned three-dimensional model retrieval method based on Laplace features and attention are completed.
[0010] Compared with the prior art, the beneficial effects of the present disclosure are as follows: In the retrieval method of the present disclosure, the Laplace-Beltrami eigenfunction can effectively capture the global geometric structure of the three-dimensional model, and has stronger stability and expressiveness than the traditional vertex space method. The model spectrum is represented by multiple perspective projections, thereby effectively alleviating the problem of loss of three-dimensional structure information in the traditional method during the view generation process, effectively avoiding the loss of structure information caused by pose, rotation, or occlusion when directly using three-dimensional geometric information, and thus improving the integrity and discriminability of the model representation; In the feature extraction stage, considering the spatial continuity of the projection views of adjacent perspectives, through a view fusion method based on the attention mechanism, the geometric information of adjacent angles can be fused when each view is extracted, enhancing the context coupling between views and improving the continuity and discriminability of the feature distribution.
[0011] Adopting a view grouping strategy effectively makes up for the defects of traditional pooling operations in dealing with detailed features, effectively retains local significant features, improves the problem of weakened expression in the scenario of a model with rich details, and significantly improves the representation ability and retrieval accuracy of the model descriptor.
[0012] The advantages of the present disclosure and the advantages of additional aspects will be described in detail in the following specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The accompanying drawings forming a part of the present disclosure are used to provide a further understanding of the present disclosure. The schematic embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute a limitation to the present disclosure.
[0014] Figure 1 is a flowchart of the retrieval method of Embodiment 1 of the present disclosure; Figure 2 is a schematic structural diagram of the model feature extraction network of Embodiment 1 of the present disclosure; Figure 3 is a schematic flowchart of the model feature extraction network implementing retrieval in Embodiment 1 of the present disclosure; DETAILED DESCRIPTION OF THE EMBODIMENTS The present disclosure will be further described below in conjunction with the accompanying drawings and embodiments.
[0015] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which the present disclosure belongs.
[0016] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof. It should be noted that, without conflict, the various embodiments and features in the present disclosure can be combined with each other. The embodiments will be described in detail below with reference to the drawings.
[0017] Embodiment 1 In the technical solutions disclosed in one or more embodiments, as Figures 1 to 3 shown, a three-dimensional model retrieval method based on Laplace features and attention includes the following steps: Step 1: Use the Laplace-Beltrami eigenfunction to represent the three-dimensional model to be retrieved; Step 2: Render the low-frequency Laplace-Beltrami eigenfunction obtained by expressing the three-dimensional model from multiple perspectives to generate multiple sets of two-dimensional projection views; Step 3: Extract features from the two-dimensional projection views, and adaptively fuse the feature information of the current view and adjacent views based on the attention mechanism to obtain the multi-perspective features after extraction; Step 4: Group the multi-perspective features, fuse the view features within the group to form group-level features, aggregate the group-level features to form the model descriptor of the final three-dimensional model, and match the closest three-dimensional model based on the model descriptor; The retrieval method of this embodiment can effectively capture the global geometric structure of a 3D model by using the Laplace - Beltrami eigenfunction, and has stronger stability and expressiveness compared with the traditional vertex space method. It represents the model spectrum using multiple perspective projections, thus effectively alleviating the problem of 3D structure information loss in the traditional method during the view generation process, and effectively avoiding the loss of structure information caused by pose, rotation, or occlusion when directly using 3D geometric information, thereby enhancing the integrity and discriminability of the model representation; in the feature extraction stage, considering the spatial continuity of the projection views of adjacent perspectives, through a view fusion method based on the attention mechanism, the geometric information of adjacent angles can be fused when each view is extracted, enhancing the context coupling between views and improving the continuity and discriminative ability of the feature distribution. At the same time, a view grouping strategy is adopted to effectively make up for the defects of traditional pooling operations in dealing with detailed features, effectively retain local significant features, improve the problem of weakened expression in the scenario of a model with rich details, and significantly improve the representation ability and retrieval accuracy of the model descriptor.
[0018] Specifically, in this embodiment, the 3D model to be retrieved is the query model; In step 1, the method of using the Laplace - Beltrami eigenfunction to express the 3D model to be retrieved, that is, calculating the Laplace - Beltrami eigenfunction of the input 3D model to be retrieved, includes the following steps: Step 11: Construct the eigenfunction Mathematically express the surface of the 3D model to be retrieved and calculate the function of the Laplace - Beltrami operator; Specifically, the function type of the eigenfunction can be a smooth function, a second - order differentiable function, a first - order differentiable function, etc. Preferably, in this embodiment, a second - order differentiable function is adopted. Assume that the eigenfunction is a second - order differentiable function defined on the surface of the input 3D model to be retrieved, and the Laplace - Beltrami operator of the eigenfunction is defined as follows: (1); Among them, represents the gradient of the eigenfunction , represents the divergence of the gradient .
[0019] Step 12: Use the Laplace - Beltrami operator to operate on the eigenfunction Perform operations to obtain the spectral features of the 3D model to be retrieved; that is, perform eigenanalysis on the Laplace-Beltrami operator Δ of the 3D model to be retrieved to obtain eigenvalues and Laplace-Beltrami eigenfunctions, thereby constituting spectral features; Specifically, construct and solve the following Helmholtz equation to obtain the Laplace-Beltrami eigenvalues of the 3D model to be retrieved and Laplace-Beltrami eigenfunctions : (2); The different Laplace-Beltrami eigenvalues obtained by solving correspond to different frequency modes, and the eigenfunctions represent the distribution of these frequency modes on the model surface; these constitute the spectral features of the 3D model. The larger the eigenvalue , the higher the frequency of the eigenfunction, and the more detailed features of the 3D model are captured.
[0020] In the above solution, to address the problem of loss of spatial structure information during the projection of the 3D model, the Laplace-Beltrami eigenfunction is used to represent the 3D model. This eigenfunction has isometric transformation invariance, can characterize the features of the 3D model, and eigenfunctions of different frequencies can capture features at different scale levels of the model. At the same time, the low-frequency eigenfunction has strong robustness to noise. The low-frequency Laplace-Beltrami function is used to retain the geometric structure information of the 3D model and strengthen the geometric representation; Based on the above representation results of the 3D model, in this embodiment, the low-frequency Laplace-Beltrami eigenfunctions of the 3D model are rendered from multiple perspectives to generate multiple sets of two-dimensional projection views, realizing the description of the non-rigid 3D model, thereby effectively alleviating the problem of loss of model structure information in traditional view generation methods.
[0021] Step 2: Generation of eigenfunction projection views: Render the low-frequency Laplace-Beltrami eigenfunctions obtained from the representation of the 3D model from multiple perspectives to generate multiple sets of two-dimensional projection views, including the following steps: Step 21: Select the first low-frequency eigenfunctions from the low-frequency Laplace-Beltrami eigenfunctions obtained from the representation of the 3D model; Step 22: Render and project each low-frequency eigenfunction from multiple preset perspectives to generate two-dimensional projection views; specifically, render and project the selected low-frequency eigenfunctions from perspectives for each eigenfunction to generate two-dimensional projection views; Step 23: For each perspective, perform channel concatenation on the corresponding two-dimensional projection view of the features to obtain the overall view representation under the corresponding perspective; wherein, and are both set values; Specifically, calculate the projection views of the first low-frequency Laplace-Beltrami eigenfunctions. Based on the Laplace-Beltrami eigenfunctions obtained in Step 1, select the first low-frequency eigenfunctions, and render and project each eigenfunction from perspectives to generate two-dimensional projection views. Denote the projection view of the th Laplace-Beltrami eigenfunction under the th perspective as , wherein , . For each perspective , perform channel concatenation on the projection views of the corresponding m eigenfunctions to obtain the overall view representation under this perspective.
[0022] Step 3: Feature aggregation of adjacent views based on attention: Extract features from the two-dimensional projection views, and adaptively fuse the feature information of the current view and adjacent views based on the attention mechanism to obtain the multi-perspective features after extraction, including the following steps: Step 31: For the overall view representation obtained from the two-dimensional projection views under each perspective, use a convolutional neural network (CNN) to extract features to obtain a view feature set , wherein, represents the view feature of the i th perspective; Step 32: Take each view feature as the original view feature in turn, and extract the view features of the adjacent perspectives of each original view feature to construct an adjacent view group to model the continuity and semantic relationship between views.
[0023] Specifically, for each original view feature , select the view features of its two adjacent perspectives, denoted as and respectively, where represents the number of perspectives, represents the modulo operation. , and together constitute the adjacent view group .
[0024] Step 33: Adopt the attention mechanism to perform information interaction and feature fusion within the adjacent view group to obtain the features after attention interaction; Specifically, the method for generating the features after attention fusion includes the following steps: Step 331: Perform transformation on the adjacent view group to generate the feature representation of attention, including query vector ( ), key vector ( ), and value vector ( ); For each adjacent group, introduce the attention mechanism to achieve information interaction and feature enhancement between views. Taking the th adjacent view group as an example , first use the fully connected layer to perform transformation on respectively to generate three new feature representations: query vector ( ), key vector ( ), and value vector ( ), and the specific representation is as follows: (3); Among them, , are learnable weight matrices respectively.
[0025] Step 332: Calculate the attention weights based on the obtained feature representation of attention to obtain the attention weight matrix between views within the adjacent group; Specifically, first calculate the dot product of the query vector ( ) and the key vector ( ), then introduce the scaling factor and use the function to normalize the attention scores to obtain the attention weight matrix between views within the adjacent group.
[0026] Step 333: Use the attention weight matrix to perform weighted summation on the value vector ( ) to obtain the fused feature , thus completing the information interaction and fusion between view features: (4); Step 334: Use a multi-layer perceptron (MLP) to perform view feature aggregation processing on the fused feature to obtain the new view feature ; Step 34: For the features after attention fusion and the original view features Concatenate and fuse through a multi-layer perceptron to obtain the view feature set of all viewpoints, which is used as the multi-viewpoint feature after the attention mechanism processes; Specifically, is concatenated with the initial view feature and further through to obtain the view feature . The view features of all viewpoints form a new feature set
[0027] Regarding the problem that the currently commonly used method of independently extracting the view features of each view fails to fully capture the context correlation and complementary information between views, which limits the expression ability of the overall model descriptor. Considering that there is significant spatial continuity and semantic correlation between adjacent viewpoints. The above solution in this embodiment realizes feature fusion and association between views based on the attention mechanism to mine the potential relationships between views, constructs adjacent view groups to enhance feature representation, thereby enhancing the robustness and discriminability of the 3D model representation, and further improving the stability and accuracy of 3D model retrieval.
[0028] In this embodiment, by introducing the attention mechanism, the feature information of the current view and adjacent views is adaptively fused, thereby enhancing the discriminability of the single-view feature, effectively mining the potential information association of the multi-viewpoint views, and overcoming the problem of insufficient view relationship modeling in the traditional method.
[0029] In step 4, to further improve the global expression ability of the model descriptor, this embodiment introduces a grouping aggregation strategy. Group the multi-viewpoint features, fuse the view features within the group to form group-level features. Aggregate the group-level features to form the model descriptor of the final 3D model. This strategy effectively improves the precision and robustness of 3D model retrieval while enhancing the model feature expression ability.
[0030] In step 4, the method for generating the model descriptor of the final 3D model includes the following steps: Step 41: Fuse the multi-viewpoint features through a fully connected process, and construct a quantization function to quantify the discrimination degree of the views; View feature fusion based on weight grouping. Input the view feature into the fully connected layer to obtain a set of scalars , and use the quantization function to quantify the discrimination degree of the views, which is defined as follows: (5); where, represents the absolute value function, is the logarithmic function, The function is used to map the result to Interval.
[0031] Step 42: Calculate the view scores based on the constructed quantization function, and group the two-dimensional projection views of multiple perspectives based on the view scores. Specifically, take as the view score, and according to divide the views into groups, which are respectively denoted as , ,…, .
[0032] Step 43: For the view features within the same group, perform average pooling operation to obtain the feature descriptor of each group, that is, the within-group feature. For the view features within the same group, perform average pooling operation to obtain group-level feature descriptors, which are respectively denoted as , ,…, .
[0033] Step 44: Calculate the average value of the view scores within each group to obtain the weighted weight for each group. To quantify the importance of each view group, let represent the number of views in the th group of views . The average value of the view scores of the th group of views is calculated as follows: The calculation formula is as follows: (6); Step 45: Weight the feature descriptors of each group based on the weighted weight to obtain the model descriptor of the three-dimensional model. Specifically, perform weighted summation on group-level feature descriptors to obtain the final model feature descriptor : (7); In Step 4, match the most similar three-dimensional model based on the model descriptor, and adopt a similarity measurement method. Specifically, judge the correlation of the feature descriptors of the three-dimensional model based on the Euclidean distance, and obtain the most similar three-dimensional model based on the correlation: For the similarity measurement of three-dimensional models, specifically: Let , respectively represent the feature descriptors of two three-dimensional models, and the calculation formula of the Euclidean distance is as follows: (8); The greater the Euclidean distance, the smaller the similarity; the smaller the Euclidean distance, the greater the similarity. The 3D model with a large similarity is the retrieval structure. Sort the 3D models in the model library and the query model in ascending order of the Euclidean distance, and return the 3D model with the smallest Euclidean distance as the retrieval result.
[0034] The above steps can be implemented by constructing a model feature extraction network. The model feature extraction network, as Figure 2 shown, includes: A Laplace-Beltrami feature calculation module, configured to use the Laplace-Beltrami feature function to express the 3D model to be retrieved; A feature function projection view generation module, configured to render the low-frequency Laplace-Beltrami feature function obtained by expressing the 3D model from multiple perspectives to generate multiple groups of two-dimensional projection views; Attention-based adjacent view feature aggregation, configured to extract features from the two-dimensional projection views and adaptively fuse the feature information of the current view and the adjacent views based on the attention mechanism to obtain the multi-view features after extraction; A view feature grouping and aggregation module, configured to group the multi-view features, fuse the view features within the group to form group-level features, aggregate the group-level features to form the model descriptor of the final 3D model, and match the closest 3D model based on the model descriptor.
[0035] In this embodiment, for the rough view fusion strategy, in the feature fusion stage, simple average pooling or splicing operations are generally used, which is difficult to model the importance difference of different views in space, resulting in problems such as view redundancy and unbalanced feature expression. The above solution in this embodiment adopts a group-level feature fusion strategy based on weight grouping, and forms the final model descriptor through grouping and weighted aggregation, further improving the discriminability and robustness of the model descriptor. It effectively makes up for the defects of traditional pooling operations in processing detailed features, and can further improve the expression ability and discrimination ability of the shape descriptor.
[0036] The model descriptor of the 3D model output by the model feature extraction network in this embodiment is the fused feature vector, as Figure 3As shown, in the actual retrieval process, for the model to be queried, that is, the three-dimensional model to be retrieved, a model descriptor (i.e., feature vector) of the three-dimensional model to be retrieved is obtained through the model feature extraction network. Based on this model descriptor, the three-dimensional model can be classified to obtain the type of the model. For the three-dimensional model database, model descriptors (i.e., feature vectors) of the three-dimensional models in the database are obtained through the model feature extraction network to construct a feature vector database. The feature vectors in the feature vector database are calculated for similarity with the model descriptor of the three-dimensional model to be retrieved to obtain the most matching retrieval result.
[0037] Embodiment 2 Based on Embodiment 1, in this embodiment, a three-dimensional model retrieval system based on Laplace features and attention is provided, including: A Laplace-Beltrami feature calculation module configured to express the three-dimensional model to be retrieved using the Laplace-Beltrami feature function; A feature function projection view generation module configured to render the low-frequency Laplace-Beltrami feature function obtained by expressing the three-dimensional model from multiple perspectives to generate multiple groups of two-dimensional projection views; Attention-based adjacent view feature aggregation configured to extract features from the two-dimensional projection views and adaptively fuse the feature information of the current view and adjacent views based on the attention mechanism to obtain the extracted multi-perspective features; A view feature grouping and aggregation module configured to group the multi-perspective features, fuse the view features within the group to form group-level features, and aggregate the group-level features to form the model descriptor of the final three-dimensional model, and match the closest three-dimensional model based on the model descriptor.
[0038] It should be noted here that each module in this embodiment corresponds one by one to each step in Embodiment 1, and the specific implementation process is the same, so it will not be repeated here.
[0039] Embodiment 3 Taking the application of the three-dimensional model in the industrial design field as an example, in this embodiment, for the specific retrieval process, the three-dimensional model is a three-dimensional model of an industrial product, which can be a three-dimensional model of a machine, equipment or component, a three-dimensional model of an actual component, etc. The three-dimensional model is drawn according to the actual component, and a similar three-dimensional model is queried in the library to speed up the product design; for example, in vehicle design, it can be a three-dimensional model of the vehicle's exterior, a three-dimensional model of the engine, three-dimensional models of the components of the vehicle frame assembly, etc. The retrieval process is as follows: First, perform standardized preprocessing on the input 3D model to be retrieved, unify its scale, orientation, and position, and ensure that the 3D model is aligned to a consistent coordinate system. Then, based on the view-based method, render multiple 2D projection views from this 3D model according to a fixed or adaptive camera view strategy. The number of views is generally 12, covering the main appearance features of the model.
[0040] Next, perform feature extraction on each rendered 2D view image. The extraction method is a convolutional neural network to extract image features. The features of each view extracted are subjected to information interaction through an attention mechanism, and finally grouped and fused into an overall global feature vector of the model.
[0041] In the database construction stage, perform the same processing on all product 3D models in the database in advance, extract view features and complete the fusion, and establish a 3D model feature library.
[0042] In the retrieval stage, the system receives the input of a 3D model of a target product, obtains its global feature vector after the above processing, and performs a similarity measurement of the Euclidean distance with the feature vectors in the database, and returns the Top-N most similar 3D model results.
[0043] The retrieval results can be applied to industrial design work scenarios such as product design comparison and competitor analysis, significantly improving the efficiency of design resource reuse and the automation level of patent infringement risk assessment.
[0044] Embodiment 4 This embodiment provides an electronic device, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the steps in the 3D model retrieval method based on Laplace features and attention in Embodiment 1 are completed.
[0045] Embodiment 5 This embodiment provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the steps in the 3D model retrieval method based on Laplace features and attention in Embodiment 1 are completed.
[0046] The above are only the preferred embodiments of the present disclosure and are not used to limit the present disclosure. For those skilled in the art, various changes and modifications can be made to the present disclosure. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.
[0047] Although the specific embodiments of the present disclosure have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications or deformations that can be made without creative efforts on the basis of the technical solutions of the present disclosure are still within the protection scope of the present disclosure.
Claims
1. A three-dimensional model retrieval method based on Laplacian features and attention, characterized in that It includes the following steps: Use the Laplace - Beltrami eigenfunctions to represent the 3D model to be retrieved; Render the low - frequency Laplace - Beltrami eigenfunctions obtained from representing the 3D model from multiple perspectives to generate multiple sets of 2D projection views; Extract features from the 2D projection views, and adaptively fuse the feature information of the current view and adjacent views based on the attention mechanism to obtain the multi - perspective features after extraction; Group the multi - perspective features, fuse the view features within the group to form group - level features, aggregate the group - level features to form the model descriptor of the final 3D model, and match the most similar 3D model based on the model descriptor.
2. The 3D model retrieval method based on Laplace features and attention according to claim 1, characterized in that: The method of using the Laplace - Beltrami eigenfunctions to represent the 3D model to be retrieved includes the following steps: Construct a feature function Mathematically express the surface of the three-dimensional model to be retrieved and calculate the function of the Laplace-Beltrami operator; Conduct eigen - analysis on the Laplace - Beltrami operator of the 3D model to be retrieved to obtain eigenvalues and Laplace - Beltrami eigenfunctions, constituting spectral features.
3. The three-dimensional model retrieval method based on Laplace features and attention according to claim 1, characterized in that, Render the low - frequency Laplace - Beltrami eigenfunctions obtained from representing the 3D model from multiple perspectives to generate multiple sets of 2D projection views, including the following steps: From the low-frequency Laplace-Beltrami eigenfunctions obtained from the three-dimensional model representation, select the first low-frequency eigenfunctions; Render and project each low - frequency eigenfunction from multiple preset perspectives to generate 2D projection views; For each perspective, splice the corresponding feature 2D projection views in channels to obtain the overall view representation under the corresponding perspective.
4. The 3D model retrieval method based on Laplacian features and attention according to claim 1, characterized in that, Extract features from the 2D projection views, and adaptively fuse the feature information of the current view and adjacent views based on the attention mechanism to obtain the multi - perspective features after extraction, including the following steps: For the overall view representation obtained from the 2D projection views under each perspective, use a convolutional neural network to extract features to obtain a set of view features; Take each view feature as the original view feature in turn, and extract the view features of adjacent perspectives of each original view feature to construct an adjacent view group to model the continuity and semantic relationship between views; Adopt the attention mechanism to conduct information interaction and feature fusion within the adjacent view group to obtain the features after attention fusion; Splice the features after attention fusion with the original view features, and fuse them through a multi - layer perceptron to obtain the set of view features of all perspectives, as the multi - perspective features processed by the attention mechanism.
5. The 3D model retrieval method based on Laplace features and attention according to claim 4, characterized in that: The method for generating the features after attention fusion includes the following steps: Transform the adjacent view group to generate the feature representation of attention, including query vector, key vector, and value vector; Calculate the attention weights based on the obtained feature representation of attention to obtain the attention weight matrix between views within the adjacent group; Use the attention weight matrix to perform weighted summation on the value vectors to obtain the fused features.
6. The 3D model retrieval method based on Laplace features and attention according to claim 1, characterized in that: The method for generating the model descriptor of the final 3D model includes the following steps: Fuse multi-view features through fully connected processing and quantify the discriminability of views by constructing a quantization function; Calculate view scores based on the constructed quantization function and group the two-dimensional projection views of multiple perspectives based on the view scores; For the view features within the same group, perform average pooling operation to obtain the feature descriptor of each group, that is, the within-group feature; Calculate the average value of the view scores within each group to obtain the weighted weight for each group; Weight the feature descriptors of each group based on the weighted weight to obtain the model descriptor of the three-dimensional model.
7. The 3D model retrieval method based on Laplace features and attention as claimed in claim 1, wherein: Match the most similar three-dimensional model based on the model descriptor. Specifically, judge the correlation between the feature descriptor of the three-dimensional model and the models in the model library based on the Euclidean distance, and obtain the most similar three-dimensional model based on the correlation.
8. A three-dimensional model retrieval system based on Laplace features and attention, characterized in that, Including: Laplace-Beltrami feature calculation module, configured to express the three-dimensional model to be retrieved using the Laplace-Beltrami feature function; Feature function projection view generation module, configured to render the low-frequency Laplace-Beltrami feature function obtained by expressing the three-dimensional model from multiple perspectives to generate multiple groups of two-dimensional projection views; Attention-based adjacent view feature aggregation, configured to extract features from the two-dimensional projection views and adaptively fuse the feature information of the current view and adjacent views based on the attention mechanism to obtain the extracted multi-view features; View feature grouping and aggregation module, configured to group the multi-view features, fuse the view features within the group to form group-level features, aggregate the group-level features to form the final model descriptor of the three-dimensional model, and match the most similar three-dimensional model based on the model descriptor.
9. An electronic device, characterized in that, Including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the steps in the three-dimensional model retrieval method based on Laplace features and attention according to any one of claims 1-7 are completed.
10. A computer-readable storage medium, characterized in that, For storing computer instructions, when the computer instructions are executed by the processor, the steps in the three-dimensional model retrieval method based on Laplace features and attention according to any one of claims 1-7 are completed.
Citation Information
Cited By
Three-dimensional model retrieval method and system based on high-order adjacency relation constraint
CN121074446A