Three-dimensional model shape recognition method, system and equipment based on multi-view fusion and medium

By adopting a multi-view fusion method in three-dimensional model shape recognition, combining multi-label learning and bidirectional GRU module, the problems of insufficient view feature extraction and neglected correlation features in MVCNN are solved, and more efficient feature fusion and recognition effects are achieved.

CN120014624APending Publication Date: 2025-05-16BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510081587.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing multi-view shape recognition algorithm MVCNN has the problem of insufficient view feature extraction and ignoring the correlation features between views in the multi-view feature fusion stage.

Method used

Using a three-dimensional model shape recognition method based on multi-view fusion, a shape recognition model including a single-view image feature extraction module, a two-way GRU module, a feature fusion network and a prediction layer is constructed, and view features are fully extracted and fused through the combination of multi-label learning and bidirectional GRU module.

Benefits of technology

It significantly improves the feature fusion effect, enhances the comprehensiveness and accuracy of feature expression, and achieves higher accuracy and robustness of three-dimensional model shape recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014624A_ABST
    Figure CN120014624A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional model shape recognition method, system and device based on multi-view fusion and a medium. The method comprises the steps that a three-dimensional model and a corresponding shape label are acquired; obtaining multi-view image data corresponding to each three-dimensional model based on a 3D rendering method; a three-dimensional model shape recognition model comprising a single-view image feature extraction module, a bidirectional GRU module, a feature fusion network and a prediction layer which are connected in sequence is constructed, the three-dimensional model shape recognition model is trained, the number of views of the model and model hyper-parameters are adjusted according to a training result, and an optimized three-dimensional model shape recognition model is obtained; and inputting a to-be-recognized three-dimensional model multi-view image into the optimized three-dimensional model shape recognition model, and outputting a shape recognition result. According to the method, the model training process is optimized while the collaboration among the views is considered, the view fusion effect is enhanced, the accuracy and robustness of three-dimensional model shape recognition are improved, and the method has high engineering application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer graphics, and in particular relates to a three-dimensional model shape recognition method, system, device and medium based on multi-view fusion. Background Art

[0002] In recent years, with the rapid development of technologies such as industrial design, virtual reality, and digital twins, the number and scale of three-dimensional models have grown rapidly in various businesses. In the field of product design and manufacturing, 3D models have become a key tool for modern industrial design with their advantages of visualization, virtualization, and digitization. It can help engineers intuitively understand the form and structure of products, thereby greatly improving design efficiency. According to surveys and statistical analysis, in the development of new products, only 20% of the designs of engineers are completely new, 40% are reused from existing models, and the remaining 40% are improvements to existing models. How to efficiently retrieve similar models based on user intent is a problem that engineers have been studying. 3D shape recognition and classification, as a basic technology for achieving shape understanding, intelligent design, and model retrieval, has long been a research hotspot in the field of computer graphics. How to accurately and reliably realize the shape recognition of three-dimensional models is of great significance to promoting the automation and intelligence of product design. However, due to the diversity and complexity of three-dimensional models, there are still many technical challenges in the shape recognition process, such as subtle differences in geometric structures between models and shape transformations under different perspectives. Therefore, innovative research on these issues is particularly important.

[0003] In recent years, deep learning technology has made great progress in the fields of image analysis and natural language processing, and has gradually been applied to the shape analysis and recognition of three-dimensional models. Compared with traditional three-dimensional model recognition methods based on manually designed features, deep learning methods can automatically extract complex and highly recognizable features through end-to-end learning, significantly improving recognition accuracy and efficiency. With the help of advanced technologies such as convolutional neural networks (CNN), graph neural networks (GNN), and Transformer, researchers have developed a variety of efficient three-dimensional model shape recognition methods, providing new technical paths for the retrieval application of three-dimensional models. These methods show great advantages in accuracy and versatility, but their application is also accompanied by challenges. For example, the generalization ability of the model needs to strike a balance between different shape categories and geometric complexity; multi-view methods need to solve the problems of view redundancy and information fusion; and high computational complexity also puts higher requirements on real-time performance and resource utilization in actual scenes. Therefore, combining domain knowledge to optimize model design, improve information fusion algorithms, and explore lightweight deep learning architectures have become important directions for future development. In the current multi-view shape recognition algorithm MVCNN, there are problems such as insufficient view feature extraction and ignoring the correlation features between views in the multi-view feature fusion stage. Summary of the invention

[0004] The purpose of the present invention is to provide a three-dimensional model shape recognition method, system, device and medium based on multi-view fusion to solve the problems existing in the above-mentioned prior art.

[0005] To achieve the above object, the present invention provides a 3D model shape recognition method based on multi-view fusion, comprising:

[0006] Acquire a three-dimensional model dataset, wherein the three-dimensional model dataset includes a plurality of three-dimensional models and corresponding shape labels;

[0007] Acquire multi-view image data corresponding to each of the three-dimensional models based on a 3D rendering method;

[0008] Constructing a three-dimensional model shape recognition model, wherein the three-dimensional model shape recognition model includes a single-view image feature extraction module, a bidirectional GRU module, a feature fusion network and a prediction layer connected in sequence;

[0009] Training the three-dimensional model shape recognition model based on the multi-view image data and the corresponding shape labels, and adjusting the number of views and model hyperparameters of the model according to the training results to obtain an optimized three-dimensional model shape recognition model;

[0010] The optimized 3D model shape recognition model is applied to the 3D model shape recognition task, multi-view images of the 3D model to be recognized are input, and the shape recognition result is output.

[0011] Optionally, the acquiring the multi-view image data corresponding to each of the three-dimensional models based on the 3D rendering method specifically includes:

[0012] Performing dimensionality reduction processing on each of the three-dimensional models, and constructing a rendering environment based on a preset blank Blender model file, importing each of the three-dimensional models after dimensionality reduction processing into the rendering environment, and scaling all the three-dimensional models to a uniform scale with the center of the three-dimensional model as the origin; achieving multi-view rendering of the model within a 360-degree range by setting evenly distributed rotation angles, and after each rotation, using Blender's camera view to capture the model image of the current view;

[0013] Each of the three-dimensional models is processed in a loop to obtain multi-view image data corresponding to each of the three-dimensional models.

[0014] Optionally, the capturing of the model image of the current viewing angle by using the camera view of Blender specifically includes:

[0015] With the center of the three-dimensional model as the origin, several virtual cameras are set up, and the model image is acquired through each virtual camera; wherein each virtual camera is aligned with the origin of the coordinate system and is set around the model, each virtual camera is evenly distributed at intervals of 30 degrees along the horizontal plane, and each camera observes the model at a depression angle of 180 degrees.

[0016] Optionally, the training of the three-dimensional model shape recognition model based on the multi-view image data and the corresponding shape labels specifically includes:

[0017] Inputting a single-view image into a single-view image feature extraction module, extracting view features, calculating a total loss function, updating the network weights of the single-view image feature extraction module by back propagation based on the total loss function, and optimizing the single-view image feature extraction module based on an Adam optimizer and a learning rate dynamic adjustment strategy to obtain a trained single-view image feature extraction module; wherein the total loss function includes a weighted combination of classification loss and consistency loss;

[0018] Input the multi-view image into the trained single-view image feature extraction module, extract each view feature and construct a view feature sequence;

[0019] The view feature sequence is input into the bidirectional GRU module to obtain a bidirectional feature representation, the bidirectional feature representation is input into the feature fusion network, global features and complementary information between views are extracted through the feature fusion network, and final fusion features are generated. The final fusion features are input into the prediction layer for classification prediction, and training is performed with the goal of minimizing the loss between the classification prediction result and the shape label corresponding to the view feature sequence, so as to obtain a trained three-dimensional model shape recognition model.

[0020] Optionally, adjusting the number of views and model hyperparameters of the model according to the training results specifically includes:

[0021] Perform performance tests on the 3D model shape recognition model under different number of view inputs, and determine the optimal number of views based on the performance test results;

[0022] The hyperparameters in the 3D model shape recognition model corresponding to the optimal number of views are dynamically adjusted to obtain an optimized 3D model shape recognition model.

[0023] A 3D model shape recognition system based on multi-view fusion, comprising:

[0024] A data acquisition module is used to obtain a three-dimensional model data set, wherein the three-dimensional model data set includes a plurality of three-dimensional models and corresponding shape labels; and obtain multi-view image data corresponding to each of the three-dimensional models based on a 3D rendering method;

[0025] A model building module, used to build a 3D model shape recognition model, wherein the 3D model shape recognition model includes a single view image feature extraction module, a bidirectional GRU module, a feature fusion network and a prediction layer connected in sequence;

[0026] The model optimization module is used to train the 3D model shape recognition model according to the multi-view image data and the corresponding shape labels, adjust the number of views and model hyperparameters of the model according to the training results, and obtain an optimized 3D model shape recognition model; apply the optimized 3D model shape recognition model to the 3D model shape recognition task, input the multi-view image of the 3D model to be recognized, and output the shape recognition result.

[0027] An electronic device comprises a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute a three-dimensional model shape recognition method based on multi-view fusion.

[0028] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the three-dimensional model shape recognition method based on multi-view fusion is implemented.

[0029] The technical effects of the present invention are:

[0030] The present invention adopts a feature extraction method of multi-label learning. Compared with other methods, it fully considers the model label characteristics of the view. The view features extracted by this method have stronger aggregation, which effectively improves the quality of subsequent feature fusion.

[0031] The present invention takes into account the correlation features between views. Compared with MVCNN, it innovatively introduces a bidirectional GRU module for feature fusion between views, which significantly improves the fusion effect and enhances the comprehensiveness and accuracy of feature expression.

[0032] The present invention verifies the model performance under various view number configurations. Compared with other methods, it achieves a sufficient balance between information coverage and computational complexity, making the trained model more suitable for practical engineering applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0034] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0035] Figure 1 A multi-view rendering flow chart in an embodiment of the present invention;

[0036] Figure 2 is a flow chart of a three-dimensional model shape recognition method in an embodiment of the present invention;

[0037] Figure 3 A schematic diagram showing a comparison of classification accuracy of different weight loss functions in an embodiment of the present invention;

[0038] Figure 4 Schematic diagram for comparing the convergence effects of loss functions in an embodiment of the present invention;

[0039] Figure 5 A schematic diagram showing comparison of classification accuracy and average category accuracy under different numbers of views in an embodiment of the present invention;

[0040] Figure 6 The figure is a flow chart of three-dimensional model shape recognition in an embodiment of the present invention. DETAILED DESCRIPTION

[0041] Various exemplary embodiments of the present invention will now be described in detail. This detailed description should not be considered as limiting the present invention, but should be understood as a more detailed description of certain aspects, features, and embodiments of the present invention.

[0042] It should be understood that the terms described in the present invention are only for describing special embodiments and are not intended to limit the present invention. In addition, for the numerical range in the present invention, it should be understood that each intermediate value between the upper and lower limits of the scope is also specifically disclosed. Each smaller range between the intermediate value in any stated value or stated range and any other stated value or intermediate value in the described range is also included in the present invention. The upper and lower limits of these smaller ranges can be independently included or excluded in the scope.

[0043] It will be apparent to those skilled in the art that various modifications and variations may be made to the specific embodiments of the present invention description without departing from the scope or spirit of the present invention. Other embodiments derived from the present invention description will be apparent to those skilled in the art. The present application description and examples are exemplary only.

[0044] The words “include,” “including,” “have,” “contain,” etc. used in this article are open-ended terms, meaning including but not limited to.

[0045] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0046] Embodiment 1

[0047] like Figure 1 - Figure 6 As shown, in this embodiment, a three-dimensional model shape recognition method based on multi-view fusion is provided, including: obtaining a three-dimensional model data set, the three-dimensional model data set including a plurality of three-dimensional models and corresponding shape labels; obtaining multi-view image data corresponding to each of the three-dimensional models based on a 3D rendering method; constructing a three-dimensional model shape recognition model, the three-dimensional model shape recognition model including a single-view image feature extraction module, a bidirectional GRU module, a feature fusion network and a prediction layer connected in sequence; training the three-dimensional model shape recognition model based on the multi-view image data and the corresponding shape labels, adjusting the number of views and model hyperparameters of the model according to the training results, to obtain an optimized three-dimensional model shape recognition model; applying the optimized three-dimensional model shape recognition model to a three-dimensional model shape recognition task, inputting a multi-view image of the three-dimensional model to be recognized, and outputting a shape recognition result.

[0048] The purpose of this embodiment is to propose a 3D model shape recognition method based on multi-view fusion. Aiming at the problems of insufficient view feature extraction and ignoring the correlation features between views in the multi-view feature fusion stage of the current multi-view shape recognition algorithm MVCNN, a new improvement scheme is proposed. In the single-view feature extraction stage, this embodiment incorporates the model label features of the view into the loss function to guide the single-view feature extraction to train in the direction of the model; in the multi-view feature fusion stage, a bidirectional gated recurrent unit (GRU) is used for bidirectional feature extraction to fully capture the correlation information from other views and further enhance the feature fusion capability; finally, the optimal number of views under this model is found through simulation experiments. The 3D model shape recognition method proposed in this embodiment optimizes the model training process while taking into account the synergy between views, enhances the view fusion effect, improves the accuracy and robustness of 3D model shape recognition, and has a high engineering application value.

[0049] The specific implementation process of this embodiment includes:

[0050] The first step is to generate multi-view image data from a three-dimensional model through 3D rendering.

[0051] In the second step, we design a multi-label loss function and feature extraction network, input a single-view image, and train the model.

[0052] The third step is to design the bidirectional GRU module and feature fusion network, input multi-view images, and train the model.

[0053] The fourth step is to adjust the number of views and model hyperparameters based on the training results to optimize the model performance.

[0054] The multi-view rendering flowchart, the 3D model shape recognition method flowchart, the classification accuracy comparison of different weight loss functions, the loss function convergence effect comparison and the loss function convergence effect comparison are as follows Figure 1 , Figure 2 , Figure 3 , Figure 4 and Figure 5 shown.

[0055] The first step is to generate multi-view image data from a three-dimensional model through 3D rendering:

[0056] 1) Construct model rendering environment:

[0057] The formats of 3D models are diverse and the structures are complex. Using 3D models directly as input objects for deep learning will significantly increase the amount of calculation and complexity. Therefore, in order to improve the computational efficiency and reduce the processing difficulty, this embodiment first performs dimensionality reduction processing on the 3D model. The 3D model dataset ModelNet used in this embodiment is stored in .off format, and the Blender software does not natively support direct processing of .off format models. Therefore, the off2obj tool is first used to convert the .off format model into .obj format, and it is imported into the Blender environment for subsequent rendering processing.

[0058] This embodiment generates multi-view images for the three-dimensional model in the Blender environment. In order to allow the python script to directly call Blender and generate multi-view images for the three-dimensional model in batches, it is necessary to first build a preset blank Blender model file and pre-configure the rendering environment such as background, texture, and lighting. In order to ensure that the model avoids the impact of surface light and shadow changes on the quality of multi-view images during the import and rendering process, thereby reducing the accuracy of subsequent recognition and retrieval, this embodiment uses the Phong lighting model (Phong Lighting Model) to complete the rendering of multiple views of the three-dimensional model. The Phong lighting model can effectively simulate the interaction between the light source and the surface of the object, and avoid complex reflection calculations between objects by defining simplified ambient light. The Phong lighting model consists of three components: ambient lighting, diffuse reflection, and specular reflection. The calculation formula is:

[0059] I=I pa k a +∑(I pa k dcosi+I ps k s cos n θ)#(1)

[0060] In the formula, k a ,k d ,k s They are respectively the ambient reflection coefficient, diffuse reflection coefficient, and specular reflection coefficient, and k d +k s = 1. According to the formula, when the color and intensity of the light source are determined, the color of the reflected light is mainly determined by the incident angle of the light source and the viewing angle.

[0061] 2) Set up the virtual camera:

[0062] Due to the diversity of models, the information contained in the multi-view images generated from different perspectives is significantly different. Therefore, when selecting the perspective, it should be ensured that the generated perspective image can show the details and features of the model as comprehensively as possible. After many experiments and performance comparison analysis, this embodiment selects the following multi-view settings:

[0063] The center of the 3D model is used as the origin of the rectangular coordinate system. All 3D models are scaled to a uniform scale, and 12 virtual cameras are set up. Each camera observes the model at a 180-degree depression angle and is evenly distributed at intervals of 30 degrees along the horizontal plane, surrounding the model. Each camera is aimed at the origin of the coordinate system. Under this setting, each 3D model can generate 12 consecutive images from different perspectives. These images are independent of each other and can capture all the information of the model from multiple perspectives.

[0064] 3) Multi-view rendering design process:

[0065] The multi-view rendering process needs to be implemented by designing a python script. First, batch load 3D model files (support .off format) from the specified path and convert them to .obj format through the off2obj tool; then, import the converted .obj file into the Blender environment and preprocess the model, including setting the geometric center as the model origin, normalizing the model size to standardize the model size, and adjusting the model position and rotation mode to ensure the consistency of the perspective. Then, by setting the evenly distributed rotation angle, the model can be rendered in multiple perspectives within 360 degrees. After each rotation, the model image of the current perspective is captured using Blender's camera view, and the generated multi-view image is saved in .png format. Finally, clean up the imported model and its related data, loop through all model files, and realize the batch generation of multi-view images of 3D models, providing a clear and effective data set for subsequent deep learning models.

[0066] The second step is to design a feature extraction network and a multi-label loss function, input a single-view image, and train the model:

[0067] The technical solution of this embodiment designs a feature extraction network and a multi-label loss function for the single-view image input of the 3D model, optimizes the feature representation and training effect, and achieves better multi-label learning. The specific design solution is as follows:

[0068] 1) Design loss function:

[0069] This embodiment uses a multi-label loss function to process the correspondence between a single-view image and multiple labels. The loss function can simultaneously optimize the feature representation at the category level (classification accuracy) and the model level (multi-view consistency). The loss function is defined as:

[0070] The first part, classification loss (category level): Cross-Entropy Loss is used to ensure that each input view can be correctly classified into its category.

[0071] The second part, consistency loss (model level): introduces correlation constraints between labels, calculates the consistency of multi-view labels, and improves the aggregation effect of features at the model level.

[0072] Total loss: A weighted combination of classification loss and consistency loss, defined as:

[0073]

[0074] Where n is the number of 3D models in the category of the model, and α and β are weight hyperparameters that control the balance between classification and consistency optimization.

[0075] 2) Feature extraction network structure:

[0076] Data preparation: Single-view images are generated by rendering 3D models, ensuring that each view corresponds to a specific category label and model label.

[0077] Training phase: Input single-view image: Input the generated single-view image into the feature extraction network to extract the high-dimensional feature vector.

[0078] Calculate loss: Calculate the total loss through the multi-label loss function and update the network weights through propagation.

[0079] Iterative Optimization: Dynamically adjust loss weights during training to balance class and model consistency goals.

[0080] Optimization method: Use Adam optimizer to update parameters, and adopt dynamic adjustment strategy for learning rate to accelerate convergence and improve model performance.

[0081] In the process of view feature extraction, this embodiment adds the model label features of the view to the loss function, guides the single view feature extraction to aggregate towards the model direction, and the extracted view features are more complete.

[0082] The third step is to design a bidirectional GRU module and feature fusion network and train the model:

[0083] This embodiment proposes a feature fusion network combined with a bidirectional GRU module to fully explore the correlation and information complementarity between multiple views of a 3D model, thereby improving feature expression capabilities and model classification performance. The specific technical solution is as follows:

[0084] 1) Design of bidirectional GRU module:

[0085] The bidirectional gated recurrent unit (GRU) is a recurrent neural network module for processing sequence data. Unlike the unidirectional GRU, the bidirectional GRU can capture both forward and backward sequence information at the same time, and is suitable for extracting deep correlations between multiple views. The module structure is as follows:

[0086] Input sequence: view feature sequence [F1, F2, ..., F n ], where F i represents the feature vector of the i-th view, and n is the number of views.

[0087] Forward GRU: processes the input sequence, captures the sequential dependencies from the first view to the last view, and outputs a sequence of hidden states

[0088] Backward GRU: processes the reverse input sequence (from the last view to the first view), captures the reverse order information, and outputs a hidden state sequence

[0089] Bidirectional fusion: Concatenate or weightedly fuse the forward and backward hidden states to obtain a bidirectional feature representation for each view:

[0090]

[0091] 2) Feature fusion network structure:

[0092] For the feature sequence [H1, H2, ..., H n ] performs pooling operation to generate the global feature vector F global , through the maximum pooling to retain key feature information:

[0093] F global =MaxPool([H1, H2, ..., H n ])#(4)

[0094] Adopting the multi-layer perceptron (MLP) structure, combined with the global feature F global And the local features H of each view i , to achieve multi-level feature fusion:

[0095] F fused =MLP(F global ,H1,H2,…,H n )#(5)

[0096] Output F fused The final feature representation of the 3D model has higher discrimination ability. Based on the fused features, the fully connected layer is connected to perform category prediction processing.

[0097] 3) Feature fusion model training process:

[0098] Feature extraction: Input the multi-view images into the feature extraction network, extract the features of each view, and construct the view feature sequence of the model.

[0099] View fusion: The view feature sequence is sent to the bidirectional GRU module to obtain the correlation between the previous and next sequences, generate a bidirectional feature representation, extract the global features and complementary information between views through the feature fusion network, and generate the final fused features.

[0100] Model training: Use the cross entropy loss function to calculate the difference between the predicted results and the true labels as the training target.

[0101] Iterative optimization: Update model parameters through the back-propagation algorithm to optimize feature extraction and fusion network performance.

[0102] In the multi-view feature fusion stage, this embodiment designs a bidirectional GRU module to perform feature fusion between views, fully considering the correlation between views, and the view fusion effect is more significant.

[0103] The fourth step is to adjust the number of views and model hyperparameters according to the training results to optimize the model effect:

[0104] The impact of the number of views on the model results is mainly reflected in two aspects: information coverage and computational complexity. The selection of the number of views needs to find a balance between fully expressing 3D geometric features and controlling redundancy and computational cost. Therefore, this embodiment tests the model performance under 9, 10, 12, and 14 views respectively. By comparing the training results, this embodiment finds that the performance of 12 views is the best.

[0105] In addition, the model performance is improved by setting and dynamically adjusting the learning rate, batch size, regularization parameters, and network structure-related parameters. The learning rate adopts a cosine annealing strategy and is dynamically reduced according to the training rounds, from rapid convergence in the early stage to fine optimization in the later stage to ensure stability. The regularization parameters are optimized through cross-validation to avoid overfitting of the model, and differentiated weight attenuation coefficients are set for different network modules. During the training process, the model loss is monitored and the performance indicators are verified, and the optimization is repeated until the model reaches the optimal performance.

[0106] This embodiment verifies the impact of different numbers of views on the performance of the model, fully considers and balances the two aspects of information coverage and computational complexity, and has good engineering application value.

[0107] The method provided in this embodiment is applied to shape recognition of three-dimensional models, which has the following advantages:

[0108] 1. This embodiment adopts a feature extraction method of multi-label learning. Compared with other methods, it fully considers the model label characteristics of the view. The view features extracted by this method have stronger aggregation, which effectively improves the quality of subsequent feature fusion.

[0109] 2. This embodiment takes into account the correlation features between views. Compared with MVCNN, it innovatively introduces a bidirectional GRU module for feature fusion between views, which significantly improves the fusion effect and enhances the comprehensiveness and accuracy of feature expression.

[0110] 3. This embodiment verifies the model performance under various view number configurations. Compared with other methods, it achieves a sufficient balance between information coverage and computational complexity, making the trained model more suitable for practical engineering applications.

[0111] A 3D model shape recognition system based on multi-view fusion, comprising:

[0112] A data acquisition module is used to obtain a three-dimensional model data set, wherein the three-dimensional model data set includes a plurality of three-dimensional models and corresponding shape labels; and obtain multi-view image data corresponding to each of the three-dimensional models based on a 3D rendering method;

[0113] A model building module, used to build a 3D model shape recognition model, wherein the 3D model shape recognition model includes a single view image feature extraction module, a bidirectional GRU module, a feature fusion network and a prediction layer connected in sequence;

[0114] The model optimization module is used to train the 3D model shape recognition model according to the multi-view image data and the corresponding shape labels, adjust the number of views and model hyperparameters of the model according to the training results, and obtain an optimized 3D model shape recognition model; apply the optimized 3D model shape recognition model to the 3D model shape recognition task, input the multi-view image of the 3D model to be recognized, and output the shape recognition result.

[0115] An electronic device comprises a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute a three-dimensional model shape recognition method based on multi-view fusion.

[0116] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the three-dimensional model shape recognition method based on multi-view fusion is implemented.

[0117] The above is only a preferred specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A 3D model shape recognition method based on multi-view fusion, characterized in that: include: Acquire a three-dimensional model dataset, wherein the three-dimensional model dataset includes a plurality of three-dimensional models and corresponding shape labels; Acquire multi-view image data corresponding to each of the three-dimensional models based on a 3D rendering method; Constructing a three-dimensional model shape recognition model, wherein the three-dimensional model shape recognition model includes a single-view image feature extraction module, a bidirectional GRU module, a feature fusion network and a prediction layer connected in sequence; Training the three-dimensional model shape recognition model based on the multi-view image data and the corresponding shape labels, and adjusting the number of views and model hyperparameters of the model according to the training results to obtain an optimized three-dimensional model shape recognition model; The optimized 3D model shape recognition model is applied to the 3D model shape recognition task, multi-view images of the 3D model to be recognized are input, and the shape recognition result is output.

2. The three-dimensional model shape recognition method based on multi-view fusion according to claim 1, characterized in that: The obtaining of multi-view image data corresponding to each of the three-dimensional models based on the 3D rendering method specifically includes: Performing dimensionality reduction processing on each of the three-dimensional models, and constructing a rendering environment based on a preset blank Blender model file, importing each of the three-dimensional models after dimensionality reduction processing into the rendering environment, and scaling all the three-dimensional models to a uniform scale with the center of the three-dimensional model as the origin; achieving multi-view rendering of the model within a 360-degree range by setting evenly distributed rotation angles, and after each rotation, using Blender's camera view to capture the model image of the current view; Each of the three-dimensional models is processed in a loop to obtain multi-view image data corresponding to each of the three-dimensional models.

3. The three-dimensional model shape recognition method based on multi-view fusion according to claim 2, characterized in that: The method of capturing the model image of the current perspective by using the camera view of Blender specifically includes: With the center of the three-dimensional model as the origin, several virtual cameras are set up, and the model image is acquired through each virtual camera; wherein each virtual camera is aligned with the origin of the coordinate system and is set around the model, each virtual camera is evenly distributed at intervals of 30 degrees along the horizontal plane, and each camera observes the model at a depression angle of 180 degrees.

4. The three-dimensional model shape recognition method based on multi-view fusion according to claim 1, characterized in that: The training of the three-dimensional model shape recognition model based on the multi-view image data and the corresponding shape labels specifically includes: Inputting a single-view image into a single-view image feature extraction module, extracting view features, calculating a total loss function, updating the network weights of the single-view image feature extraction module by back propagation based on the total loss function, and optimizing the single-view image feature extraction module based on an Adam optimizer and a learning rate dynamic adjustment strategy to obtain a trained single-view image feature extraction module; wherein the total loss function includes a weighted combination of classification loss and consistency loss; Input the multi-view image into the trained single-view image feature extraction module, extract each view feature and construct a view feature sequence; The view feature sequence is input into the bidirectional GRU module to obtain a bidirectional feature representation, the bidirectional feature representation is input into the feature fusion network, global features and complementary information between views are extracted through the feature fusion network, and final fusion features are generated. The final fusion features are input into the prediction layer for classification prediction, and training is performed with the goal of minimizing the loss between the classification prediction result and the shape label corresponding to the view feature sequence, so as to obtain a trained three-dimensional model shape recognition model.

5. The three-dimensional model shape recognition method based on multi-view fusion according to claim 1, characterized in that: The adjusting the number of views and model hyperparameters of the model according to the training results specifically includes: Perform performance tests on the 3D model shape recognition model under different number of view inputs, and determine the optimal number of views based on the performance test results; The hyperparameters in the 3D model shape recognition model corresponding to the optimal number of views are dynamically adjusted to obtain an optimized 3D model shape recognition model.

6. A 3D model shape recognition system based on multi-view fusion, characterized in that: include: A data acquisition module is used to obtain a three-dimensional model data set, wherein the three-dimensional model data set includes a plurality of three-dimensional models and corresponding shape labels; and obtain multi-view image data corresponding to each of the three-dimensional models based on a 3D rendering method; A model building module, used to build a 3D model shape recognition model, wherein the 3D model shape recognition model includes a single view image feature extraction module, a bidirectional GRU module, a feature fusion network and a prediction layer connected in sequence; The model optimization module is used to train the 3D model shape recognition model according to the multi-view image data and the corresponding shape labels, adjust the number of views and model hyperparameters of the model according to the training results, and obtain an optimized 3D model shape recognition model; apply the optimized 3D model shape recognition model to the 3D model shape recognition task, input the multi-view image of the 3D model to be recognized, and output the shape recognition result.

7. An electronic device, characterized in that: It comprises a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute a three-dimensional model shape recognition method based on multi-view fusion according to any one of claims 1-5.

8. A computer-readable storage medium, characterized in that: The computer program is stored therein, and when the computer program is executed by a processor, a three-dimensional model shape recognition method based on multi-view fusion as described in any one of claims 1 to 5 is implemented.