Fine-grained three-dimensional model classification method and system based on dynamic prototype learning
By employing a dynamic prototype learning method and utilizing multi-view rendering and feature optimization techniques, the problems of class imbalance and feature entanglement in fine-grained 3D classification were solved, achieving high-precision and robust fine-grained classification.
Patent Information
- Application Number
- CN202511298403.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-11
- Publication Date
- 2025-12-09
AI Technical Summary
Existing fine-grained 3D classification methods struggle to accurately identify subtle differences when faced with small inter-class variations and class imbalance. Furthermore, the black-box nature of existing models makes them unsuitable for meeting the requirements of feature discriminativeness and robustness in real-world tasks.
A dynamic prototype learning-based approach is adopted, which generates a sequence of two-dimensional view images through multi-view orthogonal projection rendering, extracts features using a shared weight encoder, constructs a shared prototype pool and performs dynamic soft allocation through optimal transfer theory, and optimizes the feature space by combining cross-entropy and contrastive loss to achieve probabilistic decision-making for fine-grained categories.
It significantly improves the model's classification accuracy and generalization performance under class imbalance conditions, enhances the ability to represent and discriminate tail classes, provides good interpretability and geometric perception capabilities, and improves the reliability of recognition in complex scenes.
Smart Images

Figure CN121095679A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of three-dimensional geometric analysis technology, and in particular relates to a fine-grained three-dimensional model classification method and system based on dynamic prototype learning. Background Technology
[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0003] With the widespread adoption of LiDAR, structured light scanning, and multi-view reconstruction technologies, the accuracy of acquiring 3D point cloud and mesh data has been significantly improved. Fine-grained 3D classification has also demonstrated significant application value in key areas such as high-fidelity digital asset classification in virtual reality, scene understanding in autonomous driving, and defect detection in industrial 3D printing.
[0004] However, existing fine-grained 3D classification systems generally have some obvious drawbacks, such as: (1) Existing methods mainly design global feature extraction architecture for coarse-grained categories. For example, the 40 target features in the ModelNet40 dataset are significantly different and easy to distinguish accurately. However, in fine-grained benchmarks such as FG3D, the inter-class differences of samples such as airplanes, cars and chairs are small and there is a serious class imbalance problem (e.g. there are 2027 car samples, but only 14 tricycle samples).
[0005] (2) Existing multi-view-based classification models, such as MVCNNnew, GVCNN, SMVCNN, DAN and VSFormer, typically use a linear classification head and use the SoftMax function to map the extracted features to the class probability space. These methods are difficult to capture subtle inter-class differences due to feature entanglement and local geometric structure filtering, and have a low recognition rate for tail classes.
[0006] (3) The black-box nature of existing methods makes them difficult to adapt to the high requirements of feature discrimination and robustness in real-world tasks. Summary of the Invention
[0007] To overcome the shortcomings of the prior art, this invention provides a fine-grained 3D model classification method and system based on dynamic prototype learning, which can break through the accuracy bottleneck of fine-grained classification tasks by analyzing interpretable features and prototype mapping relationships.
[0008] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions: The first aspect of this invention provides a fine-grained 3D model classification method based on dynamic prototype learning.
[0009] A fine-grained 3D model classification method based on dynamic prototype learning includes: Perform multi-view orthogonal projection rendering on the input 3D model to generate a sequence of 2D view images containing different perspectives; The encoder based on shared weights performs feature extraction and projection operations on the image sequences of each view, and fuses multi-view features; A clustering algorithm is used to initialize a set of learnable dynamic prototypes for each fine-grained category in the multi-view features, construct a shared prototype pool and generate a prototype cost matrix; a structured metric space is constructed by calculating the cosine similarity between the multi-view features and the prototype cost matrix; and dynamic soft allocation of the multi-view features and prototypes in the shared prototype pool is performed based on the optimal transfer theory. A dynamic update strategy and joint loss optimization are sequentially executed on the established shared prototype pool; by measuring the distance between the features of the test samples and the prototypes in the shared prototype pool, probabilistic decision-making for fine-grained categories is achieved.
[0010] Furthermore, the multi-view orthogonal projection rendering adopts the dodecahedral vertex projection method, that is, the three-dimensional model is orthogonally projected and rendered through 12 spatially evenly distributed viewpoints.
[0011] Furthermore, the shared prototype pool dynamically allocates multiple learnable prototypes for each fine-grained category, and each prototype dimension is strictly aligned with the feature space, and the orthogonality between prototypes is maintained through regularization constraints.
[0012] Furthermore, the cosine similarity calculation between multi-view features and prototype cost matrices includes: modeling the process of assigning feature descriptors extracted from the viewpoints corresponding to fine-grained categories to prototypes as an entropy-regularized optimal transmission problem, that is, iteratively solving the similarity distance between features and prototypes based on the entropy-regularized optimal transmission theory.
[0013] Furthermore, the dynamic update strategy is implemented based on an exponential moving average mechanism with adjustable momentum coefficients. Specifically, firstly, the weighted average of the feature matrix under multi-view sample conditions is calculated based on the embedded features with category information output by the encoder; then, the prototype allocation result is generated based on the soft allocation matrix, and the prototype vector is adjusted online using the momentum update mechanism in conjunction with the weighted mean matrix of multi-view sample features.
[0014] Furthermore, the joint loss optimization includes: optimizing the alignment between features and prototypes based on the cross-entropy loss function and excluding prototypes of incorrect categories; and optimizing the separation between features and irrelevant prototypes based on the contrastive loss function.
[0015] Furthermore, after determining the distance between the test sample features and the prototypes in the shared prototype pool, the probabilistic decision of the fine-grained category is implemented based on the nearest neighbor criterion.
[0016] A second aspect of the present invention provides a fine-grained 3D model classification system based on dynamic prototype learning.
[0017] A fine-grained 3D model classification system based on dynamic prototype learning, comprising: The multi-view rendering module is configured to perform multi-view orthographic projection rendering on the input 3D model to generate a sequence of 2D view images containing different perspectives. The feature extraction module is configured to: perform feature extraction and projection operations on the image sequences of each view based on a shared weight encoder, and fuse multi-view features; The prototype setup module is configured to: initialize a set of learnable dynamic prototypes for each fine-grained category in the multi-view features using a clustering algorithm, build a shared prototype pool, and generate a prototype cost matrix; The optimal transmission soft allocation module is configured to: construct a structured metric space by calculating the cosine similarity between multi-view features and prototype cost matrices; and perform dynamic soft allocation of multi-view features and prototypes in the shared prototype pool based on optimal transmission theory. The prototype update module is configured to execute a dynamic update strategy on the shared prototype pool. The joint loss optimization module is configured to perform joint loss optimization on the established shared prototype pool; The reasoning module is configured to enable probabilistic decision-making for fine-grained categories by measuring the distance between test sample features and prototypes in the shared prototype pool. A third aspect of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps of a fine-grained 3D model classification method based on dynamic prototype learning as described in the first aspect of the present invention.
[0018] The fourth aspect of the present invention provides an electronic device including a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of a fine-grained 3D model classification method based on dynamic prototype learning as described in the first aspect of the present invention.
[0019] The above one or more technical solutions have the following beneficial effects: (1) This invention uses a clustering algorithm to initialize a set of learnable dynamic prototypes for each fine-grained category in the multi-view features, constructs a shared prototype pool, and generates a prototype cost matrix; it constructs a structured metric space by calculating the cosine similarity between the multi-view features and the prototype cost matrix; and it performs dynamic soft allocation of the multi-view features and prototypes in the shared prototype pool based on optimal transport theory. Compared with the prior art, this invention can effectively enhance the representation ability and discriminative power of the tail category, and significantly improve the overall classification accuracy and generalization performance of the model under class imbalance conditions.
[0020] (2) By introducing a joint optimization loss mechanism, this invention combines cross-entropy classification loss and prototype comparison loss to explicitly constrain the feature space, thereby achieving compact intra-class features and separation of inter-class features, which can better capture subtle inter-class differences and improve the ability to distinguish highly similar subclasses.
[0021] (3) By constructing an interpretable dynamic prototype mapping relationship and a structured metric space, this invention can enable the classification process to have good interpretability and geometric perception. Compared with the existing technology, it not only improves the recognition reliability of the model in complex real scenes, but also provides a higher level of robustness and adaptability for three-dimensional fine-grained classification tasks.
[0022] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0023] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0024] Figure 1 This is a flowchart of a fine-grained 3D model classification method based on dynamic prototype learning in Embodiment 1 of the present invention.
[0025] Figure 2 This is a flowchart of multi-view rendering and feature extraction in Embodiment 1 of the present invention.
[0026] Figure 3 This is a flowchart of prototype learning and optimal transmission soft allocation in Embodiment 1 of the present invention.
[0027] Figure 4 This is a flowchart of the joint loss optimization and inference process in Embodiment 1 of the present invention.
[0028] Figure 5 This is a flowchart of model training in Embodiment 1 of the present invention. Detailed Implementation
[0029] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0030] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.
[0031] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0032] Example 1 This embodiment discloses a fine-grained 3D model classification method based on dynamic prototype learning.
[0033] like Figure 1 As shown, a fine-grained 3D model classification method based on dynamic prototype learning includes: Step S1: Perform multi-view orthographic projection rendering on the input 3D model to generate a set of 2D view image sequences containing different perspectives; Step S2: The encoder based on shared weights performs feature extraction and projection operations on the image sequences of each view, and fuses the features of multiple views; Step S3: Use a clustering algorithm to initialize a set of learnable dynamic prototypes for each fine-grained category in the multi-view features, construct a shared prototype pool and generate a prototype cost matrix; construct a structured metric space by calculating the cosine similarity between the multi-view features and the prototype cost matrix; and perform dynamic soft allocation of the multi-view features and prototypes in the shared prototype pool based on the optimal transfer theory. Step S4: Execute the dynamic update strategy and joint loss optimization sequentially on the established shared prototype pool; achieve probabilistic decision-making for fine-grained categories by measuring the distance between the features of the test samples and the prototypes in the shared prototype pool.
[0034] Based on the above process, this invention constructs a dynamic prototype learning framework with geometric perception capabilities, which can overcome the accuracy bottleneck of fine-grained classification tasks by analyzing interpretable features and prototype mapping relationships. To facilitate understanding of the technical solution of this invention, the specific implementation methods of this invention will be further explained and described below.
[0035] In step S1, the input 3D model is subjected to multi-view orthogonal projection rendering to generate a set of 2D view image sequences containing different perspectives; at the same time, the multi-view sample information is expanded by methods such as random cropping to obtain an expanded training set.
[0036] The multi-view orthogonal projection rendering uses the dodecahedral vertex projection method, which uses 12 spatially evenly distributed viewpoints to orthogonally project and render the 3D model, generating an RGB multi-view set with a resolution normalized to 224×224×3, to ensure the geometric completeness of the viewpoint coverage.
[0037] like Figure 2 As shown, the following method can be used to generate an RGB multi-view set with a resolution normalized to 224×224×3, ensuring the geometric completeness of the view coverage: 1) After preprocessing, the input 3D model is translated to the origin of the world coordinate system (0, 0, 0), and the coordinates of the model vertices are normalized and scaled to ensure that the model is located within the unit sphere. At the same time, a right-handed coordinate system is used to align the model's principal axes with the coordinate axes. PCA alignment can be selected to eliminate rotation ambiguity.
[0038] 2) Deploy 12 virtual cameras on the surface of a unit sphere, with azimuth angles as follows: To achieve uniform coverage at intervals [ , Range; elevation angle (i.e., top-down view) is fixed at... Projection should be performed orthogonally to avoid perspective distortion. The projection matrix should ignore depth scaling and preserve geometric proportions. The camera should always point to the origin.
[0039] 3) The Phong lighting model is adopted, with unidirectional parallel light, ambient light intensity of 0.2, diffuse reflection coefficient of 0.8, and no specular reflection. At the same time, the graphics rendering capabilities of OpenGL and rendering engines such as DirectX are combined to realize the complete rendering process. Finally, 12 multi-view RGB images with a fixed resolution of 224×224×3 are output, and each image is normalized at the channel level to normalize its pixel values to [0, 1]. 4) For the sample data obtained from multi-view rendering, three methods are used to expand the number of view samples: random cropping (retaining 80%), color dithering (amplitude of 0.2), and random horizontal flipping (probability p set to 0.5).
[0040] After the above steps, the 3D model is orthogonally projected and rendered, and data augmentation techniques are used to obtain information-enhanced 3D multi-view samples, which provide a data foundation for subsequent prototype learning and fine-grained classification.
[0041] In step S2, the encoder VSFormer based on shared weights performs feature extraction and projection operations on the image sequences (two-dimensional view images) of each view, and fuses the multi-view features to generate a globally consistent representation.
[0042] The shared encoder employs either the MVCNNnew architecture based on CNN or the VSFormer architecture based on ViT (taking the VSFormer architecture as an example), outputting a 512-dimensional L2-normalized feature vector through a global feature aggregation layer to support multi-view feature fusion. For example... Figure 2 As shown, the following methods can be used to achieve multi-view feature encoding, extraction, and aggregation: 1) Divide each view image into 16×16 patches, thus generating a total of N=14×14 samples; then, each patch is mapped to a 768-dimensional feature vector X through linear projection, forming a token with [CLS] label as the input sequence, which can be represented as follows:
[0043] The [CLS] tag is used for classification. express A vector space of size; X represents, This represents the Nth image patch.
[0044] 2) Using VSFormer as a shared encoder, where the ViT structure adds a learnable positional encoding matrix to each token. (Used to preserve spatial information), resulting in location encoding. Then, the results are computed through multiple layers of Transformer Blocks (12 layers in total) and multi-head attention is obtained. . No. i The output of each attention head is represented as follows:
[0045] in, This represents the attention mechanism. This represents the weight matrix used to implement the query matrix transformation. This represents the weight matrix used to perform the key matrix transformation. This represents the weight matrix used to perform value matrix transformations; Q This represents a query matrix, used to retrieve the information in a sequence that is most relevant to the current token. K Represents the key matrix, which is related to... Q Matching is performed to determine which information is most relevant to the current query; V The value matrix contains the actual content of each token and is used to generate the final view features based on attention weights.
[0046] 3) After the output of multi-head attention (with 8 attention heads), the original input features are added to the attention-weighted features through residual connections. This preserves the underlying geometric information and alleviates the gradient vanishing problem; subsequently, LayerNorm performs channel-level standardization on the features. This stabilizes the training process and enhances the intra-class consistency of view features. The final extracted optimized features... It also integrates local attention focus areas with global structural context, providing discriminative representations for subsequent fine-grained classification.
[0047] 4) After obtaining the 768-dimensional view-level features after layer normalization and L2 normalization, project them into a 512-dimensional space through a learnable fully connected layer to ensure that the feature space and the prototype space have consistent geometric properties:
[0048] in, This represents the projection matrix, used to intelligently reorganize and selectively retain discriminative information in the original high-dimensional feature space; Indicates the first v The result of linear projection of the view features.
[0049] 5) The products generated in the above steps V The 512-dimensional view vector features are stacked row by row to construct a feature matrix. Each row corresponds to a depth feature representation from a specific viewpoint. Global max pooling is performed along the view dimension (i.e., the row direction of the matrix), comparing the response intensity of each feature channel in the 12 views element-wise, retaining the most discriminative feature activation values, and finally generating the aggregated global feature vector:
[0050] Global features This will be further applied to downstream feature processing and classification decision-making processes, including the final transmission process and 3D fine-grained classification tasks.
[0051] In step S3, a clustering algorithm is used to initialize a set of learnable dynamic prototypes for each fine-grained category in the multi-view features, constructing a shared prototype pool and generating a prototype cost matrix. A structured metric space is constructed by calculating the cosine similarity between the multi-view features and the prototype cost matrix. Finally, dynamic soft allocation between the multi-view features and prototypes in the shared prototype pool is performed based on optimal transport theory. Specifically, this can be achieved through the following methods: Step S3-1: Calculate the initial prototype using a clustering algorithm, and simultaneously construct an iteratively optimizeable shared prototype pool, where each learnable prototype encodes an essential visual pattern containing fine-grained categories.
[0052] The shared prototype pool uses a hash table or priority queue structure to support fast retrieval and updates. It dynamically allocates 20 learnable prototypes to each fine-grained category, with each prototype dimension strictly aligned to the feature space, and orthogonality between prototypes maintained through regularization constraints. For example... Figure 3 As shown, the prototype initialization for depicting fine-grained features of a view can be achieved using the following method: 1) Use the K-Means++ clustering algorithm to... representation vectors Divided into KFor each cluster, calculate the class-related centroid. As an initial prototype. Among them, K Dynamically adjust based on intra-class complexity (e.g., using contour coefficients).
[0053] 2) Simultaneously, use a hash table or priority queue to store the prototype set for each category. This involves constructing a dynamic prototype pool that supports rapid retrieval and updates. During the initial training phase, prototypes for all categories are initialized using the clustering method described above and stored in the global prototype pool. ,in, C This is the general category.
[0054] After the above steps, the initial prototype and dynamic prototype pool can be obtained.
[0055] Step S3-2: By calculating the cosine similarity of the feature-prototype cost matrix, a structured metric space is constructed. Simultaneously, the entropy-regularized Sinkhorn optimal transfer algorithm is combined to perform dynamic soft allocation of multi-view features to prototypes in the shared prototype pool. Specifically: For category c ,Will Feature descriptors extracted from each perspective The process of assigning features to prototypes is modeled as an entropy-regularized optimal transfer problem. Based on the entropy-regularized optimal transfer theory, the Sinkhorn-Knopp iterative algorithm is used to efficiently solve the problem of minimizing the similarity distance between features and prototypes. A sparse and smooth feature-prototype soft assignment matrix is generated iteratively through matrix exponential operations and normalization operations. The entropy regularization coefficient is... Set it to 0.1 to achieve the optimal balance between computational efficiency and allocation accuracy. For example... Figure 3 As shown, the soft allocation and one-hot encoding conversion of dynamic prototypes can be completed using the following method: 1) Sample feature matrix for batch input ,in B To determine the batch size, calculate the feature matrix and the prototype matrix. Cosine similarity, generating a similarity matrix. The similarity calculation formula is as follows:
[0056] in, For the first i The feature vector of each sample For the first j One prototype vector, Represents the L2 norm. The generated similarity matrix. A structured metric space was constructed, in which the row vectors represent the global matching pattern of the sample with all prototypes, and the column vectors reflect the fine-grained distribution differences of prototypes of the same type, thus providing an interpretable geometric basis for subsequent classification decisions.
[0057] 2) Initialize two learnable scaling vectors and Both are L2 normalized to unit vectors, and they can be efficiently computed using the Sinkhorn-Knopp iterative algorithm.
[0058] 3) Define a kernel matrix This transforms the original linear programming problem of optimal transmission into a differentiable entropy regularization form. Among these, This is the entropy regularization coefficient, used to control the smoothness of the feature-to-prototype assignment result. This is the similarity matrix from features to prototypes.
[0059] 4) In the continuous feature space, perform 3 Sinkhorn-Knopp iterations to calculate the soft assignment matrix from intermediate class features to prototypes. :
[0060]
[0061] in, This means that the row normalization factor is updated alternately in each iteration. Normalization factor This ensures that the allocation probability satisfies the double random constraint; the final soft allocation matrix is obtained. In, elements Indicates sample i Assigned to prototype j The matching probability, and satisfying the row sum constraint. (i.e., normalization of the probability distribution for each sample) and columns and constraints (i.e., a balanced distribution among prototypes).
[0062] 5) To enhance the determinism of feature-prototype assignment, the soft assignment matrix will be used. The parameter is converted to an approximate one-heat code using a differentiable Gumbel-Softmax reparameterization technique, while the temperature parameter is set. The hard-assigned parameter is set to True. The Gumbel-Softmax transformation formula is expressed as follows:
[0063] in, This is noise from independently and identically distributed samples, used to introduce randomness into the data distribution. When hard=True, for... Perform the argmax operation, select the prototype with the highest probability in each row, and generate a strict one-hot encoding matrix. ,satisfy That is, to represent a sample i It can always be assigned the corresponding prototype. j This step ensures that each sample is explicitly associated with a prototype, facilitating subsequent prototype updates and classification processes.
[0064] In step S4, a dynamic update strategy and joint loss optimization are sequentially executed on the established shared prototype pool; by measuring the distance between the features of the test samples and the prototypes in the shared prototype pool, probabilistic decision-making for fine-grained categories is achieved. Specifically, this can be implemented using the following methods: Step S4-1: Based on online clustering and the momentum update strategy of exponential moving average (EMA), the dynamic representation optimization and distribution alignment of prototypes in the shared prototype pool are realized.
[0065] The prototype update module employs an adjustable momentum coefficient exponential moving average (EMA) mechanism to ensure the stability of each prototype's evolution. Specifically, the following method can be used to update the dynamic prototype: 1) Based on the embedded features with category information output by the encoder Calculate the weighted average of the feature matrix under multi-view sample conditions. :
[0066] in, Representing multi-view features Assigned to the first through online clustering process k A set of multi-view features for a prototype; This represents the total number of views for category c, reflecting the multi-view completeness of samples in that category.
[0067] 2) Based on the soft allocation matrix Generate prototype allocation results (i.e., the first) Class 1 (a set of prototype vectors), and a weighted mean matrix based on the features of multiple view samples. The prototype vector is adjusted online using a momentum update mechanism.
[0068] Among them, momentum coefficient Time decay strategy (Initial value) attenuation rate This ensures the stability of the model's initial convergence. The update mechanism requires two conditions to be met simultaneously: firstly, the sample classification accuracy... Secondly, the prediction confidence level is greater than 0.75. L2 normalization constraints are applied to the prototype after each update.
[0069] This step ensures the stability of each prototype's evolution.
[0070] Step S4-2: By jointly optimizing the cross-entropy classification loss and the prototype contrast loss, the inter-class separability and intra-class compactness of the prototype space are maximized while minimizing the class prediction error.
[0071] like Figure 4 As shown, the technical details of jointly optimizing the loss function are as follows: 1) To reduce the dispersion of intra-class features, this invention minimizes each L2-normalized view feature. Instead of allocating prototype sets Based on the distance, construct a compact feature space and design a multi-view intra-class feature alignment mechanism:
[0072] Where d(.) represents the cosine distance function from the view feature vector to the nearest prototype in the prototype set of its class. Therefore, a multi-view classification cross-entropy loss can be designed to simultaneously optimize the following two objectives: firstly, to enhance the correct class. c The first step is to achieve feature-prototype alignment by ensuring consistency between the features and the prototype space; the second step is to eliminate incorrect category prototypes through distance metric learning. The specific form of the cross-entropy loss function is as follows:
[0073] in, This represents each possible target category (including the correct category and all incorrect categories); It represents the complete set of all categories and is a fixed set.
[0074] 2) In order to achieve accurate alignment between multi-view embedded features and corresponding prototypes, and maximize the separation from irrelevant prototypes, this invention adopts a variant of InfoNCE and designs a view-prototype contrastive learning strategy, which simultaneously achieves the following two objectives: (1) bringing embedded features closer together Its assigned prototype (2) Increase the distance between them to achieve positive sample attraction; (3) Push away the embedded features With negative prototype set The distance is used to exclude negative samples. The contrastive loss function is defined as follows:
[0075] Among them, temperature coefficient τ Adaptive adjustment strategy ( t (For the current iteration number), dynamically adjust the steepness of the similarity distribution; from We select the perturbation prototypes with the highest Top-K similarity to achieve effective mining of hard negative samples; This represents a single negative prototype, which needs to be compared with the current features during the model learning process. They are mutually exclusive.
[0076] Through the above steps, fine-grained feature optimization can be achieved; simultaneously, the 3D model training process (such as...) Figure 5 As shown, the cross-entropy loss and contrastive loss are monitored in real time. If the joint loss does not decrease after three consecutive iterations, an early stopping strategy is used to terminate training. Otherwise, training continues until the set number of iterations is reached.
[0077] 3) The total loss of the fine-grained 3D classification model based on dynamic prototype learning in this invention integrates the cross-entropy classification loss of S61 and the contrastive loss of S62, and its mathematical form is as follows:
[0078] Among them, the weight learning strategy α A course-based learning strategy is adopted, with the value gradually increasing linearly from 0.1 to 1.0. Jointly optimizing the loss function can implicitly achieve the "high cohesion-low coupling" characteristics of the feature space.
[0079] Step S4-3: In the inference phase, by measuring the distance between the features of the test samples and the prototypes in the shared prototype pool, the probabilistic decision-making and optimization of fine-grained categories are realized based on the nearest neighbor criterion.
[0080] The classification reasoning process achieves fine-grained decision-making for 3D multi-view applications through similarity measurement and nearest neighbor learning mechanisms. Specifically, the final multi-view decision can be achieved using the following methods. Figure 3 D Fine-grained classification: 1) Calculate the L2 normalized embedding features of the test samples. With all learned category prototypes Cosine similarity matrix:
[0081] Both the features and the prototype are constrained within a unit hypersphere space. ).
[0082] 2) Employing hierarchical nearest neighbor decision-making, the predicted category is determined through two-level minimization operations, achieving end-to-end fine-grained classification inference and performance optimization:
[0083] After the above steps, for each category c ,choose K The result with the highest similarity among the prototypes (1- S (•) Minimize the similarity of all classes to achieve intra-class decision-making; at the same time, by comparing the optimal similarity of all classes, the final class can be determined to achieve inter-class decision-making.
[0084] To further demonstrate the significant effects of this invention, this embodiment conducts comprehensive performance verification on the 3D object classification benchmark ModelNet40 and the fine-grained 3D dataset FG3D. The evaluation is quantified using two core metrics: (1) Average Instance Accuracy (AIA), reflecting the model's overall classification performance across all test samples; and (2) Average Class Accuracy (ACA), eliminating the impact of class imbalance and accurately assessing the model's ability to identify low-frequency classes. All experiments use 5-fold cross-validation to calculate the confidence interval of the metrics, and simultaneously output AIA and ACA. This invention's method can be used as a plug-and-play module, inserted into five internationally leading 3D multi-view classification models: MVCNNnew, GVCNN, SMVCNN, DAN, and VSFormer. A 3D classification comparison experiment is conducted on the fine-grained dataset FG3D and the coarse-grained dataset ModelNet40. See Table 1 for details. Table 1. Classification results of multiple models for different categories.
[0085] Experiments show that, at both the coarse-grained and fine-grained levels, this method outperforms these cutting-edge 3D multi-view classification methods on the vast majority of subclasses, achieving the best classification performance. This demonstrates the rationality and effectiveness of the dynamic prototype learning-based 3D fine-grained classification method proposed in this invention.
[0086] Based on the above design, this invention innovatively constructs an interpretable fine-grained feature space through a prototype learning mechanism, using optimal transport theory from a prototype perspective to improve the accurate understanding of 3D multi-view information. Simultaneously, through online update algorithms and joint loss function design, it effectively alleviates the long-tail distribution problem in fine-grained classification tasks, achieving adaptive aggregation of multi-view features. Utilizing the complementarity between viewpoints further enhances feature representation capabilities and classification performance, providing more accurate and reliable entity classification technology for fields such as high-fidelity digital asset classification in virtual reality, autonomous driving scene understanding, and defect detection in industrial 3D printing. It has broad prospects in practical applications and significant commercial value.
[0087] Example 2 This embodiment discloses a fine-grained 3D model classification system based on dynamic prototype learning.
[0088] A fine-grained 3D model classification system based on dynamic prototype learning, comprising: The multi-view rendering module is configured to perform multi-view orthographic projection rendering on the input 3D model to generate a sequence of 2D view images containing different perspectives. The feature extraction module is configured to: perform feature extraction and projection operations on the image sequences of each view based on a shared weight encoder, and fuse multi-view features; The prototype setup module is configured to: initialize a set of learnable dynamic prototypes for each fine-grained category in the multi-view features using a clustering algorithm, build a shared prototype pool, and generate a prototype cost matrix; The optimal transmission soft allocation module is configured to: construct a structured metric space by calculating the cosine similarity between multi-view features and prototype cost matrices; and perform dynamic soft allocation of multi-view features and prototypes in the shared prototype pool based on optimal transmission theory. The prototype update module is configured to execute a dynamic update strategy on the shared prototype pool. The joint loss optimization module is configured to perform joint loss optimization on the established shared prototype pool; The reasoning module is configured to enable probabilistic decision-making for fine-grained categories by measuring the distance between test sample features and prototypes in the shared prototype pool. Example 3 The purpose of this embodiment is to provide a computer-readable storage medium.
[0089] A computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of a fine-grained 3D model classification method based on dynamic prototype learning as described in Embodiment 1 of this disclosure.
[0090] Example 4 The purpose of this embodiment is to provide an electronic device.
[0091] An electronic device includes a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps in a fine-grained 3D model classification method based on dynamic prototype learning as described in Embodiment 1 of this disclosure.
[0092] The steps and methods involved in the apparatuses of Embodiments 2, 3, and 4 above correspond to those in Embodiment 1. For specific implementation details, please refer to the relevant description section of Embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and enabling the processor to perform any of the methods in this invention.
[0093] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computer device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. The present invention is not limited to any particular combination of hardware and software.
[0094] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A fine-grained 3D model classification method based on dynamic prototype learning, characterized in that, include: Perform multi-view orthogonal projection rendering on the input 3D model to generate a sequence of 2D view images containing different perspectives; The encoder based on shared weights performs feature extraction and projection operations on the image sequences of each view, and fuses multi-view features; A clustering algorithm is used to initialize a set of learnable dynamic prototypes for each fine-grained category in the multi-view features, construct a shared prototype pool and generate a prototype cost matrix; a structured metric space is constructed by calculating the cosine similarity between the multi-view features and the prototype cost matrix; and dynamic soft allocation of the multi-view features and prototypes in the shared prototype pool is performed based on the optimal transfer theory. A dynamic update strategy and joint loss optimization are sequentially executed on the established shared prototype pool; by measuring the distance between the features of the test samples and the prototypes in the shared prototype pool, probabilistic decision-making for fine-grained categories is achieved.
2. The fine-grained 3D model classification method based on dynamic prototype learning as described in claim 1, characterized in that, The multi-view orthogonal projection rendering adopts the dodecahedral vertex projection method, that is, the three-dimensional model is orthogonally projected and rendered through 12 spatially evenly distributed viewpoints.
3. The fine-grained 3D model classification method based on dynamic prototype learning as described in claim 1, characterized in that, The shared prototype pool dynamically allocates multiple learnable prototypes for each fine-grained category, and each prototype dimension is strictly aligned with the feature space, and the orthogonality between prototypes is maintained through regularization constraints.
4. The fine-grained 3D model classification method based on dynamic prototype learning as described in claim 1, characterized in that, The calculation of cosine similarity between multi-view features and prototype cost matrices includes: modeling the process of assigning feature descriptors extracted from the viewpoints corresponding to fine-grained categories to prototypes as an entropy-regularized optimal transmission problem, that is, iteratively solving the similarity distance between features and prototypes based on the entropy-regularized optimal transmission theory.
5. The fine-grained 3D model classification method based on dynamic prototype learning as described in claim 1, characterized in that, The dynamic update strategy is based on an exponential moving average mechanism with adjustable momentum coefficients. Specifically, firstly, the weighted average of the feature matrix under multi-view sample conditions is calculated based on the embedded features with category information output by the encoder. Then, the prototype allocation result is generated based on the soft allocation matrix, and the prototype vector is adjusted online using the momentum update mechanism in conjunction with the weighted mean matrix of multi-view sample features.
6. The fine-grained 3D model classification method based on dynamic prototype learning as described in claim 1, characterized in that, The joint loss optimization includes: optimizing the alignment between features and prototypes based on the cross-entropy loss function and excluding prototypes of incorrect categories; and optimizing the separation between features and irrelevant prototypes based on the contrastive loss function.
7. The fine-grained 3D model classification method based on dynamic prototype learning as described in claim 1, characterized in that, After determining the distance between the test sample features and the prototypes in the shared prototype pool, the probabilistic decision of the fine-grained category is implemented based on the nearest neighbor criterion.
8. A fine-grained 3D model classification system based on dynamic prototype learning, characterized in that, include: The multi-view rendering module is configured to perform multi-view orthographic projection rendering on the input 3D model to generate a sequence of 2D view images containing different perspectives. The feature extraction module is configured to: perform feature extraction and projection operations on the image sequences of each view based on a shared weight encoder, and fuse multi-view features; The prototype setup module is configured to: initialize a set of learnable dynamic prototypes for each fine-grained category in the multi-view features using a clustering algorithm, build a shared prototype pool, and generate a prototype cost matrix; The optimal transmission soft allocation module is configured to: construct a structured metric space by calculating the cosine similarity between multi-view features and prototype cost matrices; and perform dynamic soft allocation of multi-view features and prototypes in the shared prototype pool based on optimal transmission theory. The prototype update module is configured to execute a dynamic update strategy on the shared prototype pool. The joint loss optimization module is configured to perform joint loss optimization on the established shared prototype pool; The reasoning module is configured to enable probabilistic decision-making for fine-grained categories by measuring the distance between test sample features and prototypes in the shared prototype pool.
9. A computer-readable storage medium having a program stored thereon, characterized in that, When executed by the processor, the program implements the steps in the fine-grained 3D model classification method based on dynamic prototype learning as described in any one of claims 1-7.
10. An electronic device, comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the fine-grained 3D model classification method based on dynamic prototype learning as described in any one of claims 1-7.
Citation Information
Cited By
Remote sensing data fusion method based on locality sensitive hashing and cross-modal attention
CN122200258A
Remote sensing data fusion method based on locality-sensitive hashing and cross-modal attention
CN122200258B