A new perspective synthesis method and device for product design images with collaborative large and small models
Through the method of collaborative size and model, a multi-dimensional sample set is constructed and expert models are trained. Combined with the multi-expert gating mechanism, the problems of uncontrollable generation of new perspective synthesis models and poor image quality are solved, achieving efficient and high-quality new perspective image synthesis.
Patent Information
- Application Number
- CN202510188685.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-20
AI Technical Summary
The existing new perspective synthesis model has problems such as uncontrollability, lack of fine-tuning strategies and poor image quality when generating new perspective images.
The method of size and model collaboration is adopted to optimize the generation of new perspective images by building a multi-dimensional sample set and training expert models, combining multi-expert gating mechanism.
Improves the quality and controllability of image synthesis in new perspectives and enhances the model's adaptability to specific types of samples.
Smart Images

Figure CN119672473B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of product conceptual design, and particularly relates to a method and device for synthesizing a new perspective of product design images through cooperation between large and small models. Background Art
[0002] In the field of product conceptual design, controllably obtaining high-quality multi-view design images from simple inputs is a task of great application significance, which can help designers comprehensively analyze and deliberate on the feasibility and deficiencies of the current design concept. In the traditional design process, designers use 3D modeling tools to model the design concept into a visualized 3D model, so as to observe the design scheme from multiple perspectives. Traditional 3D modeling methods have high requirements for the professional level of designers and greatly extend the cycle of the product from concept to prototype, which is not conducive to the rapid iteration of product concepts.
[0003] The development of large models enables artificial intelligence to have the ability to generate multi-view or 3D models from text, which can help designers transform fuzzy design ideas into high-quality images or 3D models in the early stage of design. However, existing 3D synthesis methods have deficiencies such as long synthesis time, poor 3D consistency, and difficulty in fine-grained control and editing of synthesis results, and still face difficulties in application in the field of product design.
[0004] Novel View Synthesis (NVS) studies how to synthesize images of other perspectives of the same subject through one or a small number of input images. In the product design process, applying this technology to synthesize multi-views from a single image can take into account controllability and convenience. Currently, there are mainly two technical routes in the NVS field: methods based on implicit 3D representations such as Neural Radiance Fields (NeRF) can obtain 3D information from sparse views and synthesize new views, but have problems such as slow synthesis speed and poor performance under single-view input; methods based on diffusion models use the 2D prior of text-to-image models as guidance to directly generate new views, but have problems such as poor multi-view consistency and poor adaptability to out-of-distribution data. In addition, current mainstream NVS models generally only provide full-scale training methods that require a large amount of data and training resources, and are trained through fixed loss functions, lacking means for fine-tuning specific types of samples.
[0005] Using the cooperation between large and small models to solve complex tasks is the mainstream idea in the current field of artificial intelligence. This strategy improves the generalization and accuracy of artificial intelligence in solving complex problems by using the general knowledge of large models and the domain-specific knowledge of small models.
[0006] In the aspect of large models, multi-modal large language models have demonstrated powerful capabilities in image recognition, understanding of positional relationships, etc., and can be applied to visual tasks such as combinatorial 3D generation; in the aspect of small models, the Low-Rank Adaptation (LoRA) method can highly controllably inject specific objects, styles, or concepts into the generation process in visual generation models such as text-to-image and video generation.
[0007] During the new view synthesis process, there are also similar general and domain knowledge, specifically manifested as the large model's cognition of the overall geometry and texture rationality, and the small model's specialized knowledge of specific categories or specific detail features. However, the existing new view synthesis models are uncontrollable in generation and lack fine-tuning strategies.
[0008] Therefore, there is an urgent need to propose a multi-dimensional evaluation method for new view image synthesis that coordinates large and small models to improve the quality of new view image synthesis. Summary of the Invention
[0009] The present invention provides a method for synthesizing a new view of a product design image by coordinating large and small models, which can improve the quality of new view image synthesis.
[0010] The present invention provides a method for synthesizing a new view of a product design image by coordinating large and small models, including:
[0011] Taking the product design image from one view as a sample, and the real color map and normal map from other views as labels, constructing a sample set with multiple samples. The samples include single-object samples and multi-object combination samples, and using the obtained multi-view generation model as the basic model;
[0012] Inputting the samples into the basic model, and taking the multiple single-object samples and multiple multi-object combination samples corresponding to the case where the similarity value between the synthesized view generated by the basic model and the label is lower than the similarity threshold as the first-dimensional sample set and the second-dimensional sample set respectively;
[0013] Comprehensively scoring the type, thickness, and reality of the synthesized view generated by the basic model by the large model, and taking the multiple samples corresponding to the scores lower than the set score as the third-dimensional sample set;
[0014] Randomly selecting samples from the sample set to construct a common sample set, where the common sample set does not include the samples of the first-dimensional sample set, the second-dimensional sample set, and the third-dimensional sample set;
[0015] Constructing three expert models, where the expert models are constructed by the basic model and the initial LORA. The initial LORA models of the three expert models are trained based on the first-dimensional sample set, the second-dimensional sample set, and the third-dimensional sample set respectively using different loss functions to obtain the first LORA model, the second LORA model, and the third LORA model;
[0016] Based on a common sample set, weights are assigned to the first, second, and third LORA models through a multi-expert gating mechanism to obtain a new perspective synthesis model for product design images;
[0017] During application, the product design image of one perspective is input into the new perspective synthesis model for product design images to obtain the color map and normal map of the predicted product design images under other perspectives.
[0018] Preferably, based on a common sample set, weights are assigned to the first, second, and third LORA models through a multi-expert gating mechanism to obtain a new perspective synthesis model for product design images, including:
[0019] Construct a training model, where the training model includes a frozen base model, the first, second, and third LORA models, and a trainable Router model. The Router model of the training model is trained through the common sample set to obtain the final Router model. Based on the base model, the first, second, and third LORA models, and the final Router model, a new perspective synthesis model for product design images is obtained.
[0020] Preferably, training the Router model of the training model through the common sample set to obtain the final Router model includes:
[0021] Input the samples into the base model, the first, second, and third LORA models, and the Router model respectively. The Router model assigns weights to the output results of the first, second, and third LORA models. The three output results with assigned weights are added together and then added to the output result of the base model to obtain the prediction result;
[0022] Construct a loss function based on the prediction result and the true color maps and normal maps of multiple perspectives of the product design image;
[0023] Train the Router model based on the common sample set through the loss function to obtain the final Router model.
[0024] Preferably, the Router model includes a linear mapping, an activation function, and Dropout;
[0025] Among them, after the common sample is linearly mapped multiple times, the mapping results are sequentially passed through the corresponding activation function and Dropout to obtain the weights assigned to the first, second, and third LORA models.
[0026] Preferably, the method for obtaining the similarity value between the synthetic view generated by the base model and the label includes:
[0027] The loss value between the synthetic view generated by the base model and the label is obtained through the multi-scale structural similarity loss function, and the obtained loss value is used as the similarity value between the synthetic view generated by the base model and the label.
[0028] Preferably, the multi-object combination sample includes the 3D model of the combined object, or is constructed by 3D editing and assembling a single object model.
[0029] Preferably, the large model comprehensively scores the type, thickness, and authenticity of the synthetic view generated by the base model, including:
[0030] Construct a comprehensive scoring prompt word based on four-dimensional evaluation criteria, and input the comprehensive scoring prompt word, the synthetic view generated by the base model, and the corresponding label into the large model to obtain the score of the comprehensive score;
[0031] The four-dimensional evaluation criteria include object type recognizability, geometric thickness accuracy, authenticity of the synthetic view, and overall similarity;
[0032] The object type recognizability is to judge the category of the input synthetic view and label, and compare the similarity of the judged categories;
[0033] The geometric thickness accuracy is to judge the geometric thickness consistency between the synthetic view and the label;
[0034] The authenticity of the synthetic view is to judge the structural and visual consistency between the synthetic view and the label;
[0035] The overall similarity is the probability of different perspectives of the synthetic view and the label being the same object based on object type recognizability, geometric thickness accuracy, and authenticity of the synthetic view.
[0036] Preferably, an expert model is constructed through the base model and the initial LORA, including:
[0037] The base model is the Wonder3D model, which includes a multi-view self-attention layer, a cross-domain self-attention layer, and a cross-attention layer. The initial LORA is injected into the multi-view self-attention layer, the cross-domain self-attention layer, and the cross-attention layer respectively to obtain the expert model.
[0038] Preferably, the loss function for training the LORA model , is the loss weight of the normal component of the predicted noise, is the loss weight of the color component of the predicted noise. When training the first LORA model, is taken, and when training the second LORA model, is taken, and when training the third LORA model, , , , wherein, is the normal component of the predicted noise output by the expert model, is the color component of the predicted noise output by the expert model, is the normal component of the random noise conforming to the standard Gaussian distribution, is the color component of the random noise conforming to the standard Gaussian distribution.
[0039] The present invention also provides a new perspective synthesis device for product design images with coordinated large and small models, which is characterized by including: a memory and one or more processors, wherein executable code is stored in the memory, and when the one or more processors execute the executable code, it is used to implement the new perspective synthesis method for product design images with coordinated large and small models.
[0040] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0041] The present invention can use large and small models to construct corresponding multi-dimensional sample sets for the main defects existing in new perspective image synthesis, train expert models in a single dimension through each dimension sample set, so that the trained expert models cooperate with the basic model to solve the corresponding defects, and then distribute weights to the first, second, and third LORA models through a multi-expert gating mechanism based on a common sample set without bias, so as to allocate more weights to the LORA model whose LORA dimension matches the difficulties of multi-view generation of samples more, thereby being able to efficiently and high-quality synthesize new perspective images using the method provided by the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 is a flowchart of a new perspective synthesis method for product design images with coordinated large and small models provided by a specific embodiment of the present invention;
[0043] Figure 2 is a structural diagram of training an expert model provided by a specific embodiment of the present invention;
[0044] Figure 3 is a structural diagram of training a trained model provided by a specific embodiment of the present invention;
[0045] Figure 4 is a structural diagram of the gating Router provided by a specific embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] To solve the problems in the prior art that the new view synthesis model generates uncontrollably, lacks a fine-tuning strategy, and the quality of the synthesized new view images is poor, specific embodiments of the present invention provide a large-small model collaborative multi-dimensional evaluation new view synthesis method for product conceptual design. Starting from the requirements in the field of product conceptual design and the main defects in the application of the current NVS model in this field, multiple-dimensional sample sets are selected and targeted optimization is carried out in the form of training LoRA. The method provided by the specific embodiments of the present invention, as Figure 1 shown, mainly includes three parts: S1, construction of a multi-dimensional data set through collaboration between large and small models; S2, training of a single-dimensional optimization model for each dimension; S3, training of a multi-dimensional optimization model based on a multi-expert gating mechanism.
[0047] S1. The specific embodiments of the present invention construct a multi-dimensional data set through collaboration between large and small models: The specific embodiments of the present invention select multiple evaluation dimensions required for new view synthesis for product conceptual design, and adopt a collaborative evaluation method between large and small models to construct a multi-dimensional sample set. For each selected dimension, multi-view data is rendered from multiple open-source large 3D data sets, and samples are selected using a collaborative scoring method between large and small models. In the application of other product conceptual design propositions, different targets can also be adopted under the same technical framework, and appropriate data sets can be selected for targeted optimization.
[0048] The specific embodiments of the present invention use the product design image from one view as a sample, and the true color map and normal map from other views as labels. Multiple samples are used to construct a sample set. The samples include single-object samples and multi-object combination samples, and the obtained multi-view generation model is used as the basic model.
[0049] In a specific embodiment, the data for training in the specific embodiments of the present invention is obtained from the Google Scanned Object and OmniObject3D data sets, totaling 1461 3D models. For each 3D model, in the azimuth , elevation range, the main viewing direction is randomly and uniformly sampled. Then, the blenderproc tool is used to render it to obtain the color and normal maps of six standard views.
[0050] The specific embodiments of the present invention select the following three dimensions as the evaluation angles for the NVS model for product conceptual design:
[0051] (1) Global image quality: Image quality is one of the main measurement indicators of the generation ability of the NVS model. By obtaining the loss amount through the synthetic view generated by the basic model and the true rendered image, that is, the label, the image quality is evaluated based on this loss amount. That is, by evaluating the similarity between the generated image and the label, including geometric and texture quality, fineness, etc., to characterize the overall quality of multi-view generation.
[0052] In the global image quality dimension, in the specific embodiments of the present invention, samples that the basic model is not good at processing and have poor generation quality are selected. From the aforementioned large-scale dataset, samples with a low similarity between the generation result of the basic model and the real image are selected as the first-dimension sample set by using the Multi Scale Structural Similarity Index Measure (MS-SSIM) loss function.
[0053] Specifically, in the specific embodiments of the present invention, the samples are input into the basic model, and multiple single-object samples corresponding to the case where the similarity value between the synthetic view generated by the basic model and the label is lower than the similarity threshold are used as the first-dimension sample set. This first-dimension sample set is used to train the corresponding expert model to improve the global image quality. In this embodiment, by using samples with a large similarity gap with the label as the first-dimension sample set, local optimum can be avoided and robustness can be improved during the training process.
[0054] Specifically, the method for obtaining the similarity value between the synthetic view generated by the basic model and the label includes:
[0055] The loss value between the synthetic view generated by the basic model and the label is obtained through the multi-scale structural similarity loss function, and the obtained loss value is used as the similarity value between the synthetic view generated by the basic model and the label.
[0056] (2) Combinatorial and multi-view Figure 1 Consistency: The consistency of the multi-views of the product is highly correlated with the retention of the design concept in the new view. In the NVS model based on the diffusion model, since it is necessary to make the new perspective image as similar as possible to the same input image, there may be problems with poor 3D consistency such as Multi-face among the synthesized views. At the same time, through testing, when the current NVS model processes combinatorial image inputs, due to the difficulty in accurately recognizing the positional relationship between multiple subjects, some subjects may be missing or the positional relationship may be disordered in the generated image. The expert model trained by this dimension sample set has the ability to maintain geometric consistency such as the number, shape, and positional relationship of the subjects in the synthesized new view when facing complex shapes or combinatorial multi-subject inputs.
[0057] Specifically, in the specific embodiments of the present invention, the samples are input into the basic model, and multiple multi-object combination samples corresponding to the case where the similarity value between the synthetic view generated by the basic model and the label is lower than the similarity threshold are used as the second-dimension sample set. This second-dimension sample set is used to train the corresponding expert model to improve combinatorial and multi-view Figure 1 Consistency.
[0058] In the specific embodiments of the present invention, in the dimensions of combinatorial and multi-view consistency, due to the lack of dedicated datasets and evaluation metrics for relevant dimensions, a batch of data was constructed by downloading 3D models containing combinatorial objects and assembling single models in 3D editing software. Samples representing phenomena that are difficult for pre-trained base models to handle, such as missing main parts and geometric inconsistencies, were selected from this data as the dataset for this dimension. In the above three sub-dimension datasets, the data format is true color and normal maps of six standard views, which is consistent with the format output by the pre-trained base model.
[0059] (3) Thickness and common sense matching degree: In product multi-view generation, it is important to obtain multi-view images that are consistent with the designer's cognition and conform to design common sense. However, due to the limitations of the diffusion model's capabilities, the NVS model is prone to generating images with thickness and shape inconsistent with the facts, which is particularly obvious when the information provided by the reference view is insufficient or when facing the front view input of uncommon and diverse-shaped objects. The sample set for this dimension is obtained by evaluating the authenticity of the images generated by the model through the general knowledge ability of the large model for the characteristics of common items.
[0060] In the specific embodiments of the present invention, the large model comprehensively scores the type, thickness, and authenticity of the synthetic views generated by the base model, and uses multiple corresponding samples with scores lower than the set score as the sample set for the third dimension.
[0061] Specifically, in the specific embodiments of the present invention, the large model comprehensively scores the type, thickness, and authenticity of the synthetic views generated by the base model, including:
[0062] Constructing a comprehensive scoring prompt based on four-dimensional evaluation criteria, and inputting the comprehensive scoring prompt, the synthetic view generated by the base model, and the corresponding label into the large model to obtain the score of the comprehensive evaluation;
[0063] The four-dimensional evaluation criteria include object type recognizability, geometric thickness accuracy, authenticity of the synthetic view, and overall similarity.
[0064] The object type recognizability provided by the specific embodiments of the present invention is to respectively judge the categories of the input synthetic view and label, and compare the similarity of the judged categories.
[0065] The geometric thickness accuracy provided by the specific embodiments of the present invention is to judge the consistency of geometric thickness between the synthetic view and the label.
[0066] The authenticity of the synthetic view provided by the specific embodiments of the present invention is to judge the structural and visual consistency between the synthetic view and the label.
[0067] The overall similarity provided by the specific embodiments of the present invention gives the probability that the synthetic view and the labels of different perspectives of the same object are the same based on the object type recognizability, geometric thickness accuracy, and the authenticity of the synthetic view.
[0068] Based on the above criteria, the large model can find a sample set in which the synthetic views output by the basic model and the labels have a large difference in terms of thickness, sensory authenticity, and thickness consistency. By training the expert model with this sample set, the corresponding expert model can obtain synthetic views with appropriate thickness, type, and structural authenticity from different perspectives.
[0069] In a specific embodiment, the prompts for the large language model provided in this embodiment are as follows:
[0070] Your goal is to determine whether two pictures are different perspectives of the same object. You need to make a judgment based on the following criteria:
[0071] 1. Object type: The two pictures belong to the same type of object. Please first guess which object they belong to based on these two pictures, and then compare the similarity of the two guesses.
[0072] 2. Geometric thickness: The two pictures are geometrically consistent, including size and thickness, and are consistent with the real object. Note that the pictures are taken from different angles, so after spatial rotation, special attention should be paid to the consistency of their thickness.
[0073] 3. Authenticity: Both pictures exist in reality and match the structure of the real object. Please analyze these two pictures separately and give a conclusion, which can refer to the object type you judged before.
[0074] 4. Overall similarity: Considering the above criteria comprehensively, these two pictures are different perspectives of the same object.
[0075] Before providing the answer, please carefully review these two pictures and focus on one aspect at a time. Try to judge each criterion independently. To provide the answer, please briefly analyze each of the above evaluation criteria. The analysis should be concise and accurate. For each criterion, you need to make a judgment using the following three options: 1 - Fully meets the criterion / 2 - Partially meets the criterion / 3 - Completely does not meet the criterion.
[0076] Important reminder: These pictures have a high degree of similarity. Please strictly evaluate the pictures and choose option 1 carefully. If you can distinguish the parts that do not meet the criteria through subtle details, please choose option 3.
[0077] Summarize your final decision in the last line in the format: "<Option for criterion 1><Option for criterion 2><Option for criterion 3><Option for criterion 4>".
[0078] 1. Object type: on the left, xxxx; on the right, xxxx; they are completely / partially / completely non - compliant with the standard.
[0079] 2. Geometric thickness: on the left, xxxx; on the right, xxxx; they are completely / partially / completely non - compliant with the standard.
[0080] 3. Degree of authenticity: on the left, xxxx; on the right, xxxx; they are completely / partially / completely non - compliant with the standard.
[0081] 4. Overall similarity: xxxx; they are completely / partially / completely non - compliant with the standard.
[0082] 5. Final answer: xxxx (e.g., 1 2 1 1 / 3 3 1 3 / 2 1 3 2).
[0083] (4) In addition, a batch of non - dimensional data is randomly selected from the samples not in the above - mentioned dimensional datasets as a common sample set for training and testing the gating Router module in the multi - dimensional optimization model. Since the common sample set is not divided into the above three dimensions, during the training process, a larger weight is assigned to the more difficult LORA model.
[0084] Specifically, in the specific embodiment of the present invention, samples are randomly selected from the sample set to construct a common sample set, which does not include the samples of the first - dimension sample set, the second - dimension sample set, and the third - dimension sample set.
[0085] In a specific embodiment, finally, the proportion of various samples in the dataset constructed in this embodiment is shown in Table 1. The format of each sample is the color image and normal image rendered by the same 3D model from 6 fixed viewpoints (front, left, right, back, left - front, right - front).
[0086] Table 1 Sample distribution of each - dimension dataset
[0087]
[0088] S2. The specific embodiment of the present invention trains a single - dimension expert model based on each constructed dimension sample set: As Figure 2 shown, the present invention builds an expert model by loading a plug - and - play LoRA module on the basic model, configures different loss functions to achieve controllable model fine - tuning. For each of the above - mentioned evaluation dimensions, the specific embodiment of the present invention uses the corresponding - dimension sample set as the input to train the LoRA module of the corresponding dimension, realizing a single - dimension optimized NVS model, that is, an expert model. Expert models of different dimensions can solve the problems corresponding to that dimension, thereby obtaining the corresponding synthesized view.
[0089] Specifically, the multi-view generation model of the specific embodiment of the present invention is modified based on the open-source Wonder3D model. On the basis of the pre-trained Wonder3D model, the multi-view self-attention layer, cross-domain attention layer, and cross-attention layer of the original model are edited, and LoRA is injected into the QKV matrix of each attention layer. The updated forward formula is , in the training stage, since a certain generation ability has been possessed in the early stage, instead of adopting the two-stage generation strategy of the basic model, the LoRAs of the multi-view self-attention module and the cross-domain attention module are trained simultaneously.
[0090] For the loss function part provided by the specific embodiment of the present invention, assume is the added random noise conforming to the standard Gaussian distribution, where represents the noise of the color part, represents the noise of the normal part. The predicted noise of the diffusion model is split by channel to obtain that still includes color and normal components. Among them is the target diffusion model, is the combined graph embedding vector at the time step, which also includes color and normal components, is the input reference view, represents the fixed camera parameter embedding under K target views. Select as the normal loss, as the color loss.
[0091] In a specific embodiment, the loss function for training the LORA model, where is the loss weight of the normal component of the predicted noise, is the loss weight of the color component of the predicted noise.
[0092] The global image quality dimension provided by the specific embodiment of the present invention mainly measures the global quality of the generated color and normal maps. The two are equally important in the evaluation process. Therefore, the same weight is given when calculating the loss. When training the first LORA model, take , that is, the loss formula is , and the first LORA model is obtained after training.
[0093] The thickness and common sense matching degree dimension provided by the specific embodiment of the present invention mainly evaluates the matching degree between the synthesized view and the real object, while the color loss refers to the difference between the finally synthesized RGB image and the real image in the NVS model, which is more in line with human judgment of image differences and contains more semantic information. Therefore, color loss is used as the main loss, and when training the third LORA model, take , in one embodiment, the loss formula is , and the third LORA model is obtained after training.
[0094] In the combined and multi-view consistency dimension, the geometric consistency between multiple views is mainly evaluated. Therefore, the normal loss representing geometric quality is used as the main loss, and when training the second LORA model, take , in one embodiment, the loss formula is , and the second LORA model is obtained after training.
[0095] S3. The specific embodiment of the present invention trains the training model constructed by the trained LORA model based on the common sample set: The specific embodiment of the present invention combines multiple single LoRA models by using the mixture of experts mechanism to realize the evaluation and optimization of the NVS model in multiple parallel dimensions. By training a gating module to parse the input samples, weights are assigned to different single-dimensional LoRAs, and an NVS model with multi-dimensional optimization is realized.
[0096] In one specific embodiment, based on the common sample set, weights are assigned to the first, second, and third LORA models through the multi-expert gating mechanism to obtain a new perspective synthesis model for product design images, as Figure 3 shown, including:
[0097] Construct a training model, the training model includes a frozen base model, the first, second, and third LORA models, and a trainable Router model. The final Router model is obtained by training the Router model of the training model through the common sample set. Based on the base model, the first, second, and third LORA models, and the final Router model, a new perspective synthesis model for product design images is obtained.
[0098] Specifically, obtaining the final Router model by training the Router model of the training model through the common sample set includes:
[0099] Input the samples into the base model, the first, second, and third LORA models, and the Router model respectively. The Router model assigns weights to the output results of the first, second, and third LORA models, and the sum of the three output results with weights is added to the output result of the base model to obtain the prediction result;
[0100] Construct a loss function for the true color map and normal map based on the prediction results and multiple perspectives of the product design image;
[0101] Train the Router model through the loss function based on the common sample set to obtain the final Router model.
[0102] In a specific embodiment of the present invention, the weight assignment gating module is implemented as a Router structure. As Figure 4 shown, this Router structure performs four linear mappings and corresponding activation and Dropout operations, and then maps the input image vector to the weight vectors of three LoRAs through Linear and Softmax. After training, the Router can predict the difficulty level of the input image in different dimensions, so as to assign higher weights to the more difficult LoRAs.
[0103] Let the weight matrix of the base model in the forward process be , and three LoRA low-rank matrices are introduced for the thickness and common sense matching degree, global image quality, combination and multi-perspective consistency to correct the base model in the corresponding dimensions, and their weight matrices are respectively . Let the input vector encoding information such as the front reference view and camera parameters be , and the gating Router The weight assignment vector obtained based on the input vector is , then the forward formula of each layer in the multi-dimensional optimization module is changed to , where is the matrix concatenation operation. In the training part, a randomly selected common data set is used for training to enhance the performance of the model on general samples. The loss function reuses the loss of the single-dimensional LoRA, and the same weight is used for the color and normal map .
[0104] The present invention also provides a new perspective synthesis device for product design images with cooperation between large and small models, which is characterized in that it includes: a memory and one or more processors, and executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the new perspective synthesis method for product design images with cooperation between large and small models.
Claims
1. A new perspective synthesis method for product design images with large and small models, characterized in that: include: A product design image from one perspective is used as a sample, and real color maps and normal maps from other perspectives are used as labels. A sample set is constructed using multiple samples, wherein the samples include single object samples and multi-object combination samples, and the obtained multi-perspective generation model is used as a basic model; The samples are input into the basic model, and a plurality of single object samples and a plurality of multi-object combination samples corresponding to when the similarity value between the synthetic view generated by the basic model and the label is lower than the similarity threshold are respectively used as the first dimension sample set and the second dimension sample set; The type, thickness and reality of the synthetic views generated by the basic model are comprehensively scored by the large model, and the corresponding multiple samples with scores lower than the set score are used as the third dimension sample set; Randomly select samples from the sample set to construct a common sample set, wherein the common sample set does not include samples from the first dimension sample set, the second dimension sample set, and the third dimension sample set; Construct three expert models, wherein the expert models are constructed by the basic model and the initial LORA, and the initial LORA models of the three expert models are respectively trained based on different loss functions through the first dimension sample set, the second dimension sample set and the third dimension sample set to obtain the first LORA model, the second LORA model and the third LORA model; Based on the common sample set, a new perspective synthesis model of product design images is obtained by assigning weights to the first, second and third LORA models through a multi-expert gating mechanism; When applied, a product design image from one perspective is input into the product design image new perspective synthesis model to obtain the predicted color map and normal map of the product design image from other perspectives.
2. The new perspective synthesis method for product design images based on large and small model collaboration according to claim 1 is characterized in that: Based on the common sample set, a multi-expert gating mechanism is used to assign weights to the first, second, and third LORA models to obtain a new perspective synthesis model for product design images, including: A training model is constructed, wherein the training model includes a frozen basic model, a first, second and third LORA models, and a trainable Router model. The Router model of the training model is trained by using a common sample set to obtain a final Router model. A new perspective synthesis model of product design images is obtained based on the basic model, the first, second and third LORA models, and the final Router model.
3. The new perspective synthesis method for product design images based on large and small model collaboration according to claim 2 is characterized in that: The final Router model is obtained by training the Router model of the training model with a common sample set, including: The samples are input into the basic model, the first, second and third LORA models, and the Router model respectively. The Router model assigns weights to the output results of the first, second and third LORA models. The three output results with assigned weights are added together and then added to the output result of the basic model to obtain the prediction result. Construct a loss function based on the predicted results and the true color maps and normal maps of multiple views of the product design image; Based on the common sample set, the Router model is trained through the loss function to obtain the final Router model.
4. The new perspective synthesis method for product design images based on large and small model collaboration according to claim 3 is characterized in that: The Router model includes linear mapping, activation function and Dropout; Among them, after the common samples are subjected to multiple linear mappings, the mapping results are sequentially passed through the corresponding activation functions and Dropout to obtain the weights assigned to the first, second and third LORA models.
5. The new perspective synthesis method for product design images based on large and small model collaboration according to claim 1 is characterized in that: The method for obtaining the similarity value between the synthetic view generated by the basic model and the label includes: The loss value between the synthetic view generated by the base model and the label is obtained through the multi-scale structural similarity loss function, and the obtained loss value is used as the similarity value between the synthetic view generated by the base model and the label.
6. The new perspective synthesis method for product design images based on large and small model collaboration according to claim 1 is characterized in that: The multi-object combination sample includes a 3D model of a combined object, or is constructed by 3D editing and assembling a single object model.
7. The new perspective synthesis method for product design images based on large and small model collaboration according to claim 1 is characterized in that: The synthetic views generated by the large model are scored for type, thickness, and realism, including: Based on the four-dimensional evaluation criteria, a comprehensive scoring prompt is constructed, and the comprehensive scoring prompt, the synthetic view generated by the basic model and the corresponding label are input into the large model to obtain the comprehensive scoring score; The four-dimensional evaluation criteria include object type identifiability, geometric thickness accuracy, realism of the synthetic view, and overall similarity; The object type recognizability is to make a category judgment on the input synthetic view and label, and to make a similarity comparison on the judged categories; The geometric thickness accuracy is to determine the geometric thickness consistency of the synthetic view and the label; The authenticity of the synthetic view is to judge the consistency of structure and appearance of the synthetic view and the label; The overall similarity is the probability that the synthesized view and the label are different perspectives of the same object based on the object type identifiability, geometric thickness accuracy and the realism of the synthesized view.
8. The new perspective synthesis method for product design images based on large and small model collaboration according to claim 1 is characterized in that: Build an expert model through the basic model and initial LORA, including: The basic model is a Wonder3D model, which includes a multi-view self-attention layer, a cross-domain self-attention layer and a cross-attention layer. Initial LORA is injected into the multi-view self-attention layer, the cross-domain self-attention layer and the cross-attention layer to obtain an expert model.
9. The new perspective synthesis method for product design images based on large and small model collaboration according to claim 1 or 8, characterized in that: Loss function for training the LORA model , is the loss weight for predicting the normal component of the noise, To predict the loss weight of the color component of the noise, take , when training the second LORA model, take , when training the third LORA model, , , ,in, is the normal component of the prediction noise output by the expert model, is the color component of the predicted noise output by the expert model, is the normal component of random noise that conforms to the standard Gaussian distribution, is the color component of random noise that conforms to the standard Gaussian distribution.
10. A new perspective synthesis device for product design images with large and small models, characterized in that: include: It comprises a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, it is used to implement the new perspective synthesis method for product design images in collaboration with large and small models as described in any one of claims 1-9.
Citation Information
Patent Citations
Method for synthesizing new view of single-view transparent object based on coding and decoding network
CN113506362A
Three-dimensional model generation method and device and electronic equipment
CN116843833A