A product optimal viewpoint fast selection method and device based on multi-view and cross-modal distillation
Patent Information
- Application Number
- CN202410535340.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-30
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2044-04-30
AI Technical Summary
[0008]但是该方法需要输入数十个视图才能保证选择的准确性,这导致巨大的计算资源和时间消耗
[0048]1.本发明提出了一种基于多视图和跨模态蒸馏的产品最佳视点快速选择方法,通过采用预训练的跨模态蒸馏分支,将全局几何特征向量中的3D模拟特征算子向视图特征描述算子中的2D模拟特征蒸馏,提高最佳视点选择模型对产品的3D结构特征的感知,能够极大减少输入视图数量的同时,带来更高的选择精度;
Smart Images

Figure CN118470448B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of multi-view deep learning and cross-modal distillation technology, specifically relating to a method and apparatus for rapid selection of the best viewpoint for a product based on multi-view and cross-modal distillation. Background Technology
[0002] Optimal viewpoint selection has been a classic problem in computer graphics for many years. With the increasingly widespread application of 3D models, user needs are increasingly focused on the efficient identification and analysis of 3D models from the optimal viewpoint. Optimal viewpoint selection has applications in various fields such as computer-aided design and interior design, assisting in scene navigation, scientific visualization, object recognition, mesh simplification, and camera placement. Optimal viewpoint selection algorithms play a crucial role in improving user experience and analytical efficiency.
[0003] Viewpoint selection methods can be categorized into hand-created feature-based methods and deep learning-based methods. Early approaches primarily used hand-created feature-based methods, such as integrating various measures of weighted geometric and topological complexity to search for a limited set of potential views. These methods tended to be limited to specific types of data, had limited generalization ability, and relied heavily on strong prior information, making them insufficient for handling complex models and scenarios.
[0004] In recent years, most research has focused on deep learning methods, which, in terms of selection results, can guarantee good selection accuracy on test data. However, due to limitations in computing resources, view discreteness, and processing time, existing methods often cannot balance selection speed and quality in practical applications.
[0005] The paper "A symmetry-aware alignment method for photogrammetric 3D models" published in the journal ISPRS Journal of Photogrammetry and Remote Sensing (2023), Volume 204, pp. 184-208, proposes a 3D model alignment method based on symmetry description. This method uses a point-global symmetry descriptor (PGSD) to link the optimal viewpoint selection with global symmetry, which is particularly effective for man-made objects with vertical support bases. The output direction of the algorithm is the selected optimal viewing point.
[0006] However, this method has strict prior requirements for the object's base and main axis, thus limiting its effectiveness in handling real-world task scenarios. Furthermore, the computational process is time-consuming, making it unacceptable in applications requiring real-time response and interactivity.
[0007] The paper "Unsupervised multi-view CNN for salient view selection and 3D interestpoint detection" published in the journal International Journal of Computer Vision (2022), Volume 130, pp. 1210-1027, proposes a method for selecting salient views based on multi-view computation. This method uses multi-view data to calculate and weight multiple salient maps, and then performs uniform viewpoint sampling from subdivided octahedrons to traverse and calculate the viewpoint with the best visibility.
[0008] However, this method requires dozens of views to ensure the accuracy of the selection, which leads to huge computational resource and time consumption. At the same time, due to the limitation of sampling density, the true best viewpoint may be lost. Summary of the Invention
[0009] The purpose of this invention is to provide a method and apparatus for rapidly selecting the optimal viewing point of a product based on multi-view and cross-modal distillation. By constructing a predictive model of the optimal viewing point of the product, during the training phase, cross-modal distillation is used to supplement the predictive model with 3D structural information, thereby enabling the generation of the optimal viewing point of the product with only a small number of view inputs. In the actual testing phase, operations related to cross-modal distillation are omitted, eliminating additional computing resources and computing time, thus achieving both fast and accurate prediction of the optimal viewing point of the product.
[0010] To achieve the above-mentioned objectives, the present invention provides the following technical solution:
[0011] In a first aspect, the present invention provides a method for rapidly selecting the optimal viewing point for a product based on multi-view and cross-modal distillation, comprising the following steps:
[0012] Step 1: Process the product 3D data to obtain viewpoint label data, multi-view data and 3D point cloud data. The data processing includes annotation, rendering and sampling. Construct the optimal viewpoint selection model. The optimal viewpoint selection model includes a point cloud optimal viewpoint regression network and a multi-view optimal viewpoint regression network. Input the product 3D data into the optimal viewpoint selection model to obtain the optimal viewpoint.
[0013] Step 2: Using viewpoint label data and 3D point cloud data as the first training set, train the optimal viewpoint regression network for point cloud, and use the regression layer to learn from the first training set to obtain the global geometric feature vector.
[0014] Step 3: Using viewpoint label data and multi-view data as the second training set, train the multi-view optimal viewpoint regression network, and use the multi-view fusion strategy to perform feature aggregation on the second training set to obtain the view feature descriptor operator.
[0015] Step 4: Project the view feature descriptor to obtain the first 3D simulation feature operator. Use the global geometric feature vector as the teacher model and the first 3D simulation feature operator as the student model. Use transmembrane distillation to obtain the second 3D simulation feature operator. Reproject the second 3D simulation feature operator and superimpose it with the view feature descriptor to obtain the object-level feature. Use regression loss to propagate the optimal viewpoint based on the object-level feature.
[0016] Step 5: Repeat step 4. When both distillation loss and regression loss meet the iteration stopping condition, the optimal observation viewpoint is obtained. Based on the optimal observation viewpoint, backpropagation is performed on the optimal viewpoint selection model to obtain the trained optimal viewpoint selection model.
[0017] Furthermore, in step 1, the data processing includes annotation, rendering, and sampling, specifically:
[0018] Viewpoint label data is obtained by labeling the subjects' preferences for the best viewpoints on the product's 3D data;
[0019] The product 3D data is rendered from N predefined multi-view viewpoints to obtain multi-view data, where N≥3;
[0020] The farthest point sampling algorithm is used to sample the 3D data of the product to obtain 3D point cloud data.
[0021] As a preferred option, a set of N viewpoints is predefined according to actual needs, and the viewpoints are set as the positions of the virtual camera. The positions of the virtual camera are represented using spherical coordinates.
[0022] RT=(θ,φ,r)
[0023] Where r represents the radius of the sphere, θ represents the angle between the projection of the virtual camera onto the vertical plane of the base axis and the positive x-axis, and φ is the angle between the projection of the virtual camera onto the horizontal plane of the base axis and the positive x-axis. The intensity of ambient light, diffuse light, and specular reflection is set, and the acquired 3D product data is rendered using the virtual camera to obtain multi-view data.
[0024] Preferably, M uniformly distributed random grids are sampled from the grid surface where the product's 3D data is located, converting the grid into a point cloud set. Using the farthest point sampling algorithm, m point clouds are selected from the point cloud set as 3D point cloud data, where M and m can be adjusted according to accuracy requirements.
[0025] Furthermore, in step 1, data augmentation processing including cropping, scaling, rotation, translation, and color dithering is performed on the multi-view data; and rotation augmentation processing including random angles is performed on the 3D point cloud data.
[0026] Furthermore, in step 2, the process of training the optimal viewpoint regression network using viewpoint label data and 3D point cloud data as the first training set specifically involves:
[0027] A point cloud transformer is used to extract the point cloud structure features of 3D point cloud data;
[0028] The regression loss is used to guide the point cloud optimal viewpoint regression network to learn viewpoint label data and point cloud structural features. The backpropagation algorithm is used to update the network parameters of the point cloud optimal viewpoint regression network. The training is completed when the regression loss is minimized.
[0029] Save the first guidance parameters during the training process, which include the mean and variance of the network parameters in the point cloud optimal viewpoint regression network.
[0030] Furthermore, in step 3, the process of using viewpoint label data and multi-view data as the second training set to train the multi-view optimal viewpoint regression network specifically involves:
[0031] The multi-view data is input into multiple feature extractors with shared weights to obtain the feature description vector for each multi-view data.
[0032] The position embedding of the corresponding feature description vector is obtained based on the index and pixel coordinates of the multi-view data;
[0033] Based on a multi-view fusion strategy connected by cascaded self-attention layers, feature aggregation is performed on location embedding and viewpoint label data, and the network parameters of the multi-view optimal viewpoint regression network are updated using the backpropagation algorithm.
[0034] Save the second guidance parameters during the training process. The second guidance parameters include the mean and variance of the network parameters in the multi-view optimal viewpoint regression network.
[0035] Furthermore, in step 4, the projection of the view feature description operator to obtain the first 3D simulation feature operator, using the global geometric feature vector as the teacher model and the first 3D simulation feature operator as the student model, and employing transmembrane distillation to obtain the second 3D simulation feature operator, specifically:
[0036] Based on the first and second guidance parameters, the view feature description operator is input into the projection function to obtain the first 3D simulation feature operator;
[0037] Transmembrane distillation is employed to enable the student model to learn the 3D structural information in the teacher model, thereby obtaining the second 3D simulation feature operator.
[0038] Preferably, in step 5, backpropagation uses viewpoint selection error. When the viewpoint selection error is minimized, the best trained viewpoint selection model is obtained.
[0039] Secondly, in order to achieve the above-mentioned objectives, the present invention also provides a device for rapid selection of the best viewpoint of a product based on multi-view and cross-modal distillation, including a model parameter preparation module, a global geometric feature acquisition module, a view feature acquisition module, a cross-modal distillation module, and an iterative training module;
[0040] The model parameter preparation module is used to process the product 3D data to obtain viewpoint label data, multi-view data and 3D point cloud data. The data processing includes annotation, rendering and sampling, and constructs the optimal viewpoint selection model. The optimal viewpoint selection model includes a point cloud optimal viewpoint regression network and a multi-view optimal viewpoint regression network. The product 3D data is input into the optimal viewpoint selection model to obtain the optimal viewpoint.
[0041] The global geometric feature acquisition module is used to train the optimal viewpoint regression network of the point cloud using viewpoint label data and 3D point cloud data as the first training set, and to learn the global geometric feature vector by using the regression layer on the first training set.
[0042] The view feature acquisition module is used to train the best view regression network for multiple views using viewpoint label data and multi-view data as the second training set, and to perform feature aggregation on the second training set using a multi-view fusion strategy to obtain view feature descriptors.
[0043] The cross-modal distillation module is used to project the view feature descriptor to obtain the first 3D simulation feature operator. Using the global geometric feature vector as the teacher model and the first 3D simulation feature operator as the student model, cross-modal distillation is used to obtain the second 3D simulation feature operator. The second 3D simulation feature operator is reprojected and superimposed on the view feature descriptor to become the object-level feature. The optimal viewpoint is then propagated based on the object-level feature using regression loss.
[0044] The iterative training module is used to repeat the cross-modal distillation module. When both the distillation loss and regression loss meet the iteration stopping condition, the optimal observation viewpoint is obtained. The optimal viewpoint selection model is backpropagated based on the optimal observation viewpoint to obtain the trained optimal viewpoint selection model.
[0045] Thirdly, to achieve the above-mentioned objectives, the present invention also provides a device for rapid selection of the optimal viewpoint of a product based on multi-view and cross-modal distillation, including a memory and a processor. The memory is used to store a computer program, and the processor is used to implement the method for rapid selection of the optimal viewpoint of a product based on multi-view and cross-modal distillation provided in the first aspect of the embodiments of the present invention when the computer program is executed.
[0046] Fourthly, to achieve the above-mentioned objectives, the present invention also provides a computer-readable storage medium storing a computer program that, when used with a computer, implements the method for rapid selection of the optimal viewpoint of a product based on multi-view and cross-modal distillation provided by the present invention in the first aspect.
[0047] The beneficial effects of this invention are as follows:
[0048] 1. This invention proposes a fast method for selecting the best viewpoint of a product based on multiple views and cross-modal distillation. By using a pre-trained cross-modal distillation branch, the 3D simulated feature operator in the global geometric feature vector is distilled into the 2D simulated feature in the view feature description operator, which improves the perception of the 3D structural features of the product by the best viewpoint selection model. This can greatly reduce the number of input views while bringing higher selection accuracy.
[0049] 2. The optimal viewpoint selection method for products constructed in this invention has wide applicability. It can select the best observation viewpoint for the 3D data of any product. Since there are no prior requirements for the product itself, it can greatly reduce time consumption and increase user experience. Moreover, the model has strong generalization ability and is more suitable for animation design, virtual reality, and augmented reality scenarios that require interactivity. Attached Figure Description
[0050] Figure 1 This is a flowchart of a method for rapid selection of the best viewpoint for a product based on multi-view and cross-modal distillation, provided in an embodiment of the present invention.
[0051] Figure 2 This is a detailed flowchart illustrating the method for rapid selection of the optimal viewpoint for a product based on multi-view and cross-modal distillation, provided in an embodiment of the present invention.
[0052] Figure 3 This is a diagram of the network model training and testing structure for a method for rapid optimal viewpoint selection of products based on multi-view and cross-modal distillation, provided in an embodiment of the present invention.
[0053] Figure 4 This is a qualitative comparison diagram of the optimal viewpoint selection for actual products provided by the embodiments of the present invention, wherein (a) represents the standard viewpoint sphere, (b) represents the view rendering result based on the SOTA method, and (c) represents the view rendering result of this solution. Detailed Implementation
[0054] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and do not limit the scope of protection of the invention.
[0055] like Figure 1 As shown, the embodiment provides a method for rapid selection of the optimal viewpoint for a product based on multi-view and cross-modal distillation, including the following steps:
[0056] S110, perform data processing on the product 3D data to obtain viewpoint label data, multi-view data and 3D point cloud data. The data processing includes annotation, rendering and sampling. Construct an optimal viewpoint selection model. The optimal viewpoint selection model includes a point cloud optimal viewpoint regression network and a multi-view optimal viewpoint regression network. Input the product 3D data into the optimal viewpoint selection model to obtain the optimal viewpoint.
[0057] First, 3D product data is collected. In this embodiment, two public datasets, ShapeNet and ModelNet40, are used as the 3D product data for selecting the best viewing point. Next, data processing is performed, including the following steps:
[0058] (1) The viewpoint label data of the product 3D data is obtained by annotating the optimal viewpoint preferences of S subjects. The more subjects there are, the higher the final accuracy will be. However, considering the convergence speed of the model, S is generally taken as 12 to 50. Therefore, in this embodiment, S = 26.
[0059] (2) Define a set of N viewpoints, setting each viewpoint to the position of the virtual camera. The larger the value of N, the better. Considering model convergence, N is generally set to 3-40. In this embodiment, N = 6 is selected, which are respectively... That is, the virtual camera's position distribution RT in spherical coordinates can be expressed by the formula:
[0060] RT=(θ,φ,r)
[0061] Wherein, the radius of the sphere r = 1, θ is the angle between the projection of the virtual camera onto the vertical plane of the base axis and the positive x-axis, and φ is the angle between the projection of the virtual camera onto the horizontal plane of the base axis and the positive x-axis. The product 3D data is rendered at six defined viewpoints, resulting in six 224*224 resolution multi-view data MV_data for each product 3D data.
[0062] (3) Sample 2000 uniformly distributed random grids on the grid surface where the product 3D data is located, and convert the grids into point cloud sets. Using the farthest point sampling algorithm, select 1024 point clouds from the point cloud set as 3D point cloud data PC_data.
[0063] (4) Data augmentation processing of multi-view data and 3D point cloud data: random cropping and scaling, random rotation and translation, and color dithering are performed on multi-view data. By randomly selecting three angles on [0,2π] and applying them to the x, y, and z axes respectively, rotational enhancement processing of 3D point cloud data is achieved.
[0064] S120 uses viewpoint label data and 3D point cloud data as the first training set to train the optimal viewpoint regression network for point clouds. The regression layer learns from the first training set to obtain the global geometric feature vector.
[0065] By modifying the classification layer of the PCT network to a regression layer and freezing the network weights, the optimal viewpoint regression network for point clouds proposed in this invention is constructed. Figure 2 As shown, viewpoint label data and 3D point cloud data are used as the first training set and input into the optimal viewpoint regression network for training. The prediction results are then normalized.
[0066] For the trained point cloud optimal viewpoint regression network, the input of the last layer before output, i.e., the regression prediction layer, is saved and called the global geometric feature vector f. 3d Simultaneously, the running average μ of the network parameters of the optimal viewpoint regression network for the point cloud during training is saved. 3d and variance σ 3d , as the first guiding parameter.
[0067] S130: Using viewpoint label data and multi-view data as the second training set, train the multi-view optimal viewpoint regression network, and use the multi-view fusion strategy to aggregate features of the second training set to obtain the view feature descriptor operator.
[0068] The augmented multi-view data is input into multiple ResNet50 feature extractors with shared weights. Each feature extractor uses the first four residual blocks of the ResNet50, which contain convolutional layers, batch normalization layers, and activation layers, to obtain a 7x7 dimensional feature description vector for each multi-view data point.
[0069] The location embedding (PE) corresponding to each feature descriptor vector is obtained by using the index and pixel coordinates of multi-view data. Then, a multi-view fusion strategy consisting of three cascaded self-attention layers is used to aggregate the features of the location embedding PE to obtain the view feature descriptor f. mv Save the running average μ of the network parameters of the multi-view optimal viewpoint regression network during training.mv and variance σ mv , as the second guiding parameter.
[0070] PE i(x,y) =MLP(i,x,y)
[0071] Where i represents the index of the i-th multi-view data corresponding to the current sample. (x,y) are the pixel coordinates of the current pixel in the i-th multi-view data.
[0072] S140: Project the view feature descriptor to obtain the first 3D simulation feature operator. Using the global geometric feature vector as the teacher model and the first 3D simulation feature operator as the student model, transmembrane distillation is performed to obtain the second 3D simulation feature operator. The second 3D simulation feature operator is reprojected and superimposed on the view feature descriptor to form object-level features. Using regression loss, the optimal viewpoint is propagated based on the object-level features.
[0073] The view feature description operator contains 2D information about the product. mv The global geometric feature vector contains the 3D structural information of the product. 3d An alignment module composed of multi-layer MLP encoders is used to align 2D information f. mv It is transformed into a simulated 3D feature representation through projection, namely the first 3D simulated feature operator containing 2D information. This is to enable the first 3D simulation feature operator With 3D structural information f 3d Distribution alignment enables effective feature learning. The first 3D simulated feature operator. Expressed as a formula:
[0074]
[0075] in, μ represents the first MLP layer in the alignment module. mv and σ mv σ represents the mean and variance of the network parameters in the multi-view optimal viewpoint regression network. 3d and μ 3d This represents the mean and variance of the network parameters in the optimal viewpoint regression network for point clouds.
[0076] like Figure 3 As shown, the global geometric feature vector is used as the teacher model, and the first 3D simulation feature operator is used. As a student model, transmembrane distillation is employed to enable the first 3D simulation feature operator containing 2D information. Learn the 3D structural information f in the global geometric feature vector 3dAfter learning, the second 3D simulation feature operator is obtained. Expressed as a formula:
[0077]
[0078] in, This represents the second MLP layer in the alignment module, Δσ 3d ,Δμ 3d These are learnable parameters during the distillation training process.
[0079] At this point, the learned second 3D simulation feature operator It possesses 2D information from the view feature description operator, as well as 3D structural information from the global geometric feature vector, and its form is aligned with the distribution of 3D structural information.
[0080] Utilizing distillation loss L align Ensure 3D structural information f 3d Distill as much as possible into the first 3D simulation feature operator In the middle, the distillation loss L align ,as follows:
[0081]
[0082] Where p 3d It is the predicted viewpoint of the point cloud optimal viewpoint regression network, p mv d(·) represents the predicted viewpoint of the multi-view optimal viewpoint regression network, and d(·) represents the distance between the two.
[0083] Therefore, it is further necessary to integrate the second 3D simulation feature operator. The purpose of reprojection is to transfer the learned 3D structural information... Calibration back f mv The format is optimized to achieve smooth 2D-3D feature connections. The second 3D simulated feature operator after reprojection is expressed by the formula:
[0084]
[0085] in, It is the third MLP layer in the alignment module, Δσ mv , Δμ mv These are learnable parameters during the training process.
[0086] After reprojection with f mv Superposition yields object-level features f fin It contains all the 3D structural information and 2D information of the product. That is:
[0087]
[0088] in, This represents the concat operation of features. The purpose of overlay is to enhance the understanding of 3D structural information during the learning and regression process of multi-view models, thereby improving the accuracy of viewpoint selection.
[0089] According to the regression loss L reg Utilizing object-level features f fin Regression optimization optimal viewpoint p mv The regression loss L reg ,as follows:
[0090] L reg =1-argmin d(p mv ,l),l∈l S
[0091] L total =L reg +L align
[0092] Where, p 3d It is the predicted viewpoint of the point cloud optimal viewpoint regression network, p mv It is the predicted viewpoint of the multi-view optimal viewpoint regression network; S L represents multi-labeled labeled data consisting of S human markers; align It is the distillation loss, L reg It is regression loss, L total This represents the total loss of the optimal viewpoint selection model.
[0093] S150, repeat S140. When both distillation loss and regression loss meet the iteration stopping condition, the optimal observation viewpoint is obtained. Based on the optimal observation viewpoint, the optimal viewpoint selection model is backpropagated to obtain the trained optimal viewpoint selection model.
[0094] Repeat S140 until both distillation loss and regression loss reach the iteration stopping condition, i.e., minimized. This yields object-level features containing all 3D and 2D information. Using these object-level features, regression optimization is performed to obtain the optimal viewpoint for the product.
[0095] By backpropagating the optimal viewpoint selection model using the best observation viewpoint, a trained optimal viewpoint selection model can be obtained and used for actual testing. In this embodiment, backpropagation uses the viewpoint selection error (VSE). When the viewpoint selection error VSE is minimized, the optimal viewpoint selection model with the best viewpoint selection performance under experimental conditions is obtained.
[0096] The formula for calculating VSE is as follows:
[0097]
[0098]
[0099] Where GD (Geodesic Distance) is the great circle distance (maximum π per unit sphere), and the predicted viewpoint given by the algorithm. The value of S is the number of effective human annotators, l is the label represented by S, and p... pred It is the viewpoint result predicted by the model, W S This represents a measure of the difference between the points predicted by the algorithm and the label assigned to the s-th subject. A value close to zero indicates that the algorithm has provided a view that is close to the subject's choice. Annotations from multiple annotators represent multiple labels.
[0100] When using the trained optimal viewpoint selection model for actual reasoning, you only need to input the 3D data of the corresponding product into the model to output the best viewing point of the product under different views.
[0101] like Figure 4 The figure shows a qualitative comparison of the optimal viewpoint selection model proposed in this invention when applied to the optimal viewpoint selection of actual products. In the figure, (a) is the standard viewpoint sphere, where the black "x" corresponds to the label provided by the 26 subjects in the publicly available test benchmark, the cyan "o" represents the SOTA-selected viewpoint, and the red "o" represents the viewpoint selected by the method in this paper. (b) shows the view rendering result of the SOTA-selected viewpoint, and (c) shows the view rendering result of the method in this paper. As can be seen from the figure, the method in this paper exhibits better viewpoint display performance on most models. Specifically, when comparing within the same viewpoint sphere (a), our method is closer to the label, while previous methods showed varying degrees of deviation above and below the ideal viewpoint. The optimal viewpoint selection model proposed in this invention can complete the viewpoint selection of basic 3D products within 1-2 seconds, far faster than the 1-2 minutes of previous methods. It can predict the optimal viewpoint of the generated model more quickly and accurately, and has broad applicability.
[0102] Based on the same inventive concept, this invention also provides a device for rapid selection of the best viewpoint of a product based on multi-view and cross-modal distillation, including a model parameter preparation module, a global geometric feature acquisition module, a view feature acquisition module, a cross-modal distillation module, and an iterative training module.
[0103] The model parameter preparation module is used to process the product 3D data to obtain viewpoint label data, multi-view data and 3D point cloud data. The data processing includes annotation, rendering and sampling, and constructs the optimal viewpoint selection model. The optimal viewpoint selection model includes a point cloud optimal viewpoint regression network and a multi-view optimal viewpoint regression network. The product 3D data is input into the optimal viewpoint selection model to obtain the optimal viewpoint.
[0104] The global geometric feature acquisition module is used to train the optimal viewpoint regression network of the point cloud using viewpoint label data and 3D point cloud data as the first training set. The regression layer is used to learn the first training set to obtain the global geometric feature vector.
[0105] The view feature acquisition module is used to train the best view regression network for multiple views using viewpoint label data and multi-view data as the second training set, and to perform feature aggregation on the second training set using a multi-view fusion strategy to obtain the view feature descriptor.
[0106] The cross-modal distillation module is used to project the view feature descriptor to obtain the first 3D simulation feature operator. Using the global geometric feature vector as the teacher model and the first 3D simulation feature operator as the student model, cross-modal distillation is used to obtain the second 3D simulation feature operator. The second 3D simulation feature operator is reprojected and superimposed on the view feature descriptor to become the object-level feature. The optimal viewpoint is then propagated based on the object-level feature using regression loss.
[0107] The iterative training module is used to repeat the cross-modal distillation module. When both the distillation loss and the regression loss meet the iteration stopping condition, the optimal observation viewpoint is obtained. The optimal viewpoint selection model is backpropagated based on the optimal observation viewpoint to obtain the trained optimal viewpoint selection model.
[0108] For the product optimal viewpoint rapid selection device based on multi-view and cross-modal distillation provided in the embodiments of the present invention, since it basically corresponds to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present invention. Those skilled in the art can understand and implement this without creative effort.
[0109] Based on the same inventive concept, the embodiment also provides a device for rapid selection of the best viewpoint of a product based on multi-view and cross-modal distillation, including a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to implement the above-described method for rapid selection of the best viewpoint of a product based on multi-view and cross-modal distillation when the computer program is executed.
[0110] The product optimal viewpoint rapid selection device based on multi-view and cross-modal distillation proposed in this invention embodiment can be a device such as a computer. The device embodiment can be implemented through software, hardware, or a combination of both. Taking software implementation as an example, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions from non-volatile memory into memory and executing them. From a hardware perspective, in addition to the processor, memory, network interface, and non-volatile memory, the product optimal viewpoint rapid selection device based on multi-view and cross-modal distillation provided in this invention embodiment may also include other hardware depending on the actual functions of the device with data processing capabilities, which will not be elaborated further.
[0111] Based on the same inventive concept, the embodiment also provides a computer-readable storage medium storing a computer program that, when used by a computer, implements the above-described method for rapid selection of the best viewpoint of a product based on multiple views and cross-modal distillation.
[0112] The computer-readable storage medium can be an internal storage unit of any data processing device as described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium can also be an external storage device, such as a plug-in hard disk, smart media card (SMC), SD card, flash card, etc., equipped on the device. Furthermore, the computer-readable storage medium can include both internal storage units of any data processing device and external storage devices. The computer-readable storage medium is used to store the computer program and other programs and data required by the data processing device, and can also be used to temporarily store data that has been output or will be output.
[0113] It should be noted that the product optimal viewpoint rapid selection device, product optimal viewpoint rapid selection equipment, and computer-readable storage medium based on multi-view and cross-modal distillation provided in the above embodiments all belong to the same concept as the product optimal viewpoint rapid selection method embodiments based on multi-view and cross-modal distillation. For details of their specific implementation process, please refer to the product optimal viewpoint rapid selection method embodiments based on multi-view and cross-modal distillation, which will not be repeated here.
[0114] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention in any way. Although the implementation process of the present invention has been described in detail above, those skilled in the art can still modify the technical solutions described in the foregoing examples or make equivalent substitutions for some of the technical features. All modifications and equivalent substitutions made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for rapid selection of the optimal viewpoint for a product based on multi-view and cross-modal distillation, characterized in that, Includes the following steps: Step 1: Process the product 3D data to obtain viewpoint label data, multi-view data and 3D point cloud data. The data processing includes annotation, rendering and sampling. Construct the optimal viewpoint selection model. The optimal viewpoint selection model includes a point cloud optimal viewpoint regression network and a multi-view optimal viewpoint regression network. Input the product 3D data into the optimal viewpoint selection model to obtain the optimal viewpoint. Step 2: Using viewpoint label data and 3D point cloud data as the first training set, train the optimal viewpoint regression network for point cloud, and use the regression layer to learn from the first training set to obtain the global geometric feature vector. Step 3: Using viewpoint label data and multi-view data as the second training set, train the multi-view optimal viewpoint regression network, and use the multi-view fusion strategy to perform feature aggregation on the second training set to obtain the view feature descriptor operator. Step 4: Project the view feature descriptor to obtain the first 3D simulation feature operator. Use the global geometric feature vector as the teacher model and the first 3D simulation feature operator as the student model. Use transmembrane distillation to obtain the second 3D simulation feature operator. Reproject the second 3D simulation feature operator and superimpose it with the view feature descriptor to obtain the object-level feature. Use regression loss to propagate the optimal viewpoint based on the object-level feature. An alignment module is constructed using a multi-layer MLP encoder, and the first 3D simulation feature operator is used. Expressed as a formula: in, This indicates the first MLP layer in the alignment module. For view feature description operators, and This represents the mean and variance of the network parameters in the multi-view optimal viewpoint regression network. and This represents the mean and variance of the network parameters in the optimal viewpoint regression network for point clouds. Second 3D simulation feature operator This can be expressed as a formula: in, This indicates the second MLP layer in the alignment module. , These are learnable parameters during the distillation training process; The second 3D simulation feature operator after reprojection This can be expressed as a formula: = in, It is the third MLP layer in the alignment module. , These are learnable parameters during the training process; Step 5: Repeat step 4. When both distillation loss and regression loss meet the iteration stopping condition, the optimal observation viewpoint is obtained. Based on the optimal observation viewpoint, backpropagation is performed on the optimal viewpoint selection model to obtain the trained optimal viewpoint selection model.
2. The method of Claim 1, wherein, In step 1, the data processing includes annotation, rendering, and sampling, specifically: Viewpoint label data is obtained by labeling the subjects' preferences for the best viewpoints on the product's 3D data; The product 3D data is rendered from N predefined multi-view viewpoints to obtain multi-view data, where N≥3; The farthest point sampling algorithm is used to sample the 3D data of the product to obtain 3D point cloud data.
3. The method for rapid selection of the optimal viewpoint for a product based on multi-view and cross-modal distillation according to claim 2, characterized in that, In step 1, data augmentation processing including cropping, scaling, rotation, translation, and color dithering is performed on the multi-view data; rotation augmentation processing including random angles is performed on the 3D point cloud data.
4. The method of Claim 3, wherein, In step 2, the process of training the optimal viewpoint regression network for the point cloud using viewpoint label data and 3D point cloud data as the first training set specifically involves: A point cloud transformer is used to extract the point cloud structure features of 3D point cloud data; The regression loss is used to guide the point cloud optimal viewpoint regression network to learn viewpoint label data and point cloud structural features. The backpropagation algorithm is used to update the network parameters of the point cloud optimal viewpoint regression network. The training is completed when the regression loss is minimized. Save the first guidance parameters during the training process, which include the mean and variance of the network parameters in the point cloud optimal viewpoint regression network.
5. The method of Claim 3, wherein, In step 3, the process of using viewpoint label data and multi-view data as the second training set to train the multi-view optimal viewpoint regression network specifically involves: The multi-view data is input into multiple feature extractors with shared weights to obtain the feature description vector for each multi-view data. The position embedding of the corresponding feature description vector is obtained based on the index and pixel coordinates of the multi-view data; Based on a multi-view fusion strategy connected by cascaded self-attention layers, feature aggregation is performed on location embedding and viewpoint label data, and the network parameters of the multi-view optimal viewpoint regression network are updated using the backpropagation algorithm. Save the second guidance parameters during the training process. The second guidance parameters include the mean and variance of the network parameters in the multi-view optimal viewpoint regression network.
6. The method for rapid selection of the optimal viewpoint for a product based on multi-view and cross-modal distillation according to claim 1, characterized in that, In step 5, backpropagation uses viewpoint selection error. When the viewpoint selection error is minimized, the best trained viewpoint selection model is obtained.
7. A device for rapid selection of the optimal viewpoint for a product based on multi-view and cross-modal distillation, characterized in that, The method for rapid selection of the best viewpoint of a product based on multi-view and cross-modal distillation as described in any one of claims 1-6 includes a model parameter preparation module, a global geometric feature acquisition module, a view feature acquisition module, a cross-modal distillation module, and an iterative training module. The model parameter preparation module is used to process the product 3D data to obtain viewpoint label data, multi-view data and 3D point cloud data. The data processing includes annotation, rendering and sampling, and constructs the optimal viewpoint selection model. The optimal viewpoint selection model includes a point cloud optimal viewpoint regression network and a multi-view optimal viewpoint regression network. The product 3D data is input into the optimal viewpoint selection model to obtain the optimal viewpoint. The global geometric feature acquisition module is used to train the optimal viewpoint regression network of the point cloud using viewpoint label data and 3D point cloud data as the first training set, and to learn the global geometric feature vector by using the regression layer on the first training set. The view feature acquisition module is used to train the best view regression network for multiple views using viewpoint label data and multi-view data as the second training set, and to perform feature aggregation on the second training set using a multi-view fusion strategy to obtain the view feature descriptor. The cross-modal distillation module is used to project the view feature descriptor to obtain the first 3D simulation feature operator. Using the global geometric feature vector as the teacher model and the first 3D simulation feature operator as the student model, cross-modal distillation is used to obtain the second 3D simulation feature operator. The second 3D simulation feature operator is reprojected and superimposed on the view feature descriptor to become the object-level feature. The optimal viewpoint is then propagated based on the object-level feature using regression loss. The iterative training module is used to repeat the cross-modal distillation module. When both the distillation loss and the regression loss meet the iteration stopping condition, the optimal observation viewpoint is obtained. The optimal viewpoint selection model is backpropagated based on the optimal observation viewpoint to obtain the trained optimal viewpoint selection model.
8. A device for rapid selection of the optimal viewpoint for a product based on multi-view and cross-modal distillation, comprising a memory and a processor, the memory for storing a computer program, characterized in that, The processor is configured to, when executing the computer program, implement the method for rapid selection of the best viewpoint of a product based on multi-view and cross-modal distillation as described in any one of claims 1-6.
9. A computer-readable storage medium having stored thereon a computer program, characterized in that When the computer program is used, it implements the method for rapid selection of the best viewpoint of a product based on multi-view and cross-modal distillation as described in any one of claims 1-6.