A 3D style transfer method for single-view images
Through the texture style transfer method of double-residual gated network and Mlp network combined with semantic mask, the problem of three-dimensional deformation and texture migration is solved from a single-view picture, and 3D objects with novel shapes and textures are generated, and the diversity and richness of texture maps are improved.
Patent Information
- Application Number
- CN202211649506.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-21
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-12-21
AI Technical Summary
When the prior art expands the 2D style migration to the 3D field, it is difficult to effectively solve the problem of three-dimensional deformation starting from a single-view image. The traditional texture style migration method lacks semantic mask, resulting in insufficient richness of texture maps.
The dual-residual gated network and Mlp network are used to learn the shape characteristics of the source object and the target object through the mask mask, texture perception and key point correspondence in the two-dimensional picture as supervision signals, and the shape characteristics of the source object and the target object are fused to generate a novel 3D model. At the same time, semantic masks are introduced in traditional texture style transfer to realize partial style transfer and generate richer texture maps.
Three-dimensional deformation and texture migration starting from a single-view picture are realized, and 3D objects with novel shapes and textures are generated, which solves the problems of difficult, cost and time-consuming three-dimensional deformation in the existing technology, and improves the diversity and richness of texture maps.
Smart Images

Figure CN115775298B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of 3D model design, and particularly to a 3D style transfer method for single-view images. Background Art
[0002] Neural image style transfer has received increasing attention in the computer vision community due to its remarkable success in many automated creations. These applications mainly belong to 2D style transfer, which transfers the artistic style of a reference image to another content image. Recently, it has been extended to transfer the shape and texture style of a 3D object to another for editing 3D content in augmented reality and virtual reality. However, this extension requires obtaining the 3D information of the object, which has major problems such as high difficulty, high cost, and long time consumption.
[0003] In terms of 3D deformation, the goal of 3D deformation is to deform the shape of the source object into the target object. Some early works attempted to deform the shape of the object manually by setting control points or cages. To avoid human intervention, some works attempted to find the correspondence between the source and the target by constructing deformable cages, deformation encoding, and detected 3D key points. The above methods achieve 3D deformation by finding the correspondence between the source and the target with manual or 3D models as the input, and it is not easy to establish this correspondence.
[0004] In terms of style transfer, 2D style transfer has been widely explored. Many works have been proposed to achieve 2D style transfer of arbitrary styles, including AdaIN, LST, Adaattn, InS, and EFDM. These methods can be easily extended to UV map texture maps because these UV map texture maps are still 2D images. Following this strategy, we explored a semantic UV map texture map style transfer method to achieve the diversity of UV map textures. The purpose of 3D style transfer is to extend 2D style transfer to the 3D field to generate stylized 3D models. Most early works gave 3D models or reconstructed them from single-view or multi-view images, and then explored differential rendering methods to transfer the image style to the 3D mesh. 3DStyleNet mainly studies the shape transformation from the source 3D model to the target 3D model. In addition, many stylized 3D scene methods have been proposed to analogize 3D style transfer to 2D style transfer and perform 2D-like style transfer on point clouds or implicit fields without considering the style of the 3D shape. Summary of the Invention
[0005] The object of the present invention is to propose a three-dimensional style transfer method for single-view images. By using a double residual gated network and an Mlp network, with the mask mask, texture perception, and key point correspondence relationship in the two-dimensional picture as the supervision signals, the shape features of the source object and the target object are learned and fused to generate a novel 3D model. And a semantic mask is introduced in the traditional texture style transfer to achieve partial style transfer and generate a richer texture map.
[0006] The technical solution for implementing the present invention: In the first aspect, the present invention provides a three-dimensional style transfer method for single-view images, including the following steps:
[0007] Step 1: Given the original picture S and the target picture T of the same category, use the Encoder encoder of UMR to extract their shape features F S , F T , UV Map U S , U T and the camera pose P S , P T and the category semantic segmentation UV Map U seg ;
[0008] Step 2: Use the double residual gated network (DRG Net) to extract the shape features F S , F T the shape features that can be fused with the same semantics
[0009] Step 3: Set the fusion ratio factor α to control the degree of fusion of the shape features with the same semantics feature fusion, and fuse the shape features according to the ratio to generate a shape fusion feature
[0010] Step 4: Use the Mlp network with the shape fusion feature as the input to obtain the three-dimensional space point coordinates that can combine the shape characteristics of the two pictures and add them to the three-dimensional category common template to generate a specific three-dimensional shape model;
[0011] Step 5: Calculate the shape mask loss L mask , perceptual loss L per and three-dimensional key point correspondence loss L key , and optimize the parameters of the double residual gated network and the Mlp network;
[0012] Step 6: Use the VGG network of the trained traditional texture style transfer to extract the UV Map U S , U T texture feature V S , V T ;
[0013] Step 7: Downsample the category semantic segmentation UV Map U seg to the texture feature V S , V T scale, and respectively add the other semantic mask parts except the background part to the texture feature V seg , V S to form each semantic texture feature T
[0014] Step 8: Use the mean and variance of the semantic mask part of the semantic texture feature to replace the mean and variance of the semantic mask part of the semantic texture feature , and input it into the trained Decoder network of traditional texture style transfer to generate the semantic texture style transfer result, which together with the three-dimensional shape model generated in Step 4 constitutes an innovative three-dimensional model of shape and texture.
[0015] Preferably, in Step 1, the Encoder of UMR is used to extract the shape features F S , F T of the original picture S and the target picture T respectively as the preliminary shape features, which have rich shape information. The UV Map U S , U T and the category semantic segmentation UV Map U seg have a fixed mapping method with the three-dimensional model space points and textures and are semantically consistent, which can support the semantic invariance of the stylized UV Map generated by subsequent texture transfer.
[0016] Preferably, in the double-residual gated network in Step 2, the iterative structure is used to hierarchically extract the shape features of the original picture and the target picture, the residual structure is used to maintain the original information of the shape features, and the gated information is generated based on the shape features of both to obtain the residual change of the shape features.
[0017] Preferably, the scale factor in Step 3 controls the fusion ratio of the shape features of the original picture and the target picture.
[0018] Preferably, in Step 4, the Mlp network takes the shape fusion feature as the input to obtain the three-dimensional space point coordinates that can combine the shape characteristics of the two pictures and add them to the three-dimensional category common template to generate a specific three-dimensional shape model.
[0019] Preferably, in Step 5, the generated innovative three-dimensional model is respectively projected onto the two-dimensional plane at the original picture camera pose and the target picture camera pose, and the IOU is calculated between the mask after projection and the masks of the original picture and the target picture to generate L mask , and calculate the perceptual loss L between the texture - mapped image after projection onto the two - dimensional plane and the original image and the target image per , and calculate the key - point loss L for key points after projection key ; Through L mask and L per to control the overall shape of the 3D model, and L key can maintain the relative positions between various parts of the 3D model.
[0020] Preferably, in step 6, use the pre - trained VGG network for traditional texture style transfer to extract the UV Map U S , U T texture feature V S , V T .
[0021] Preferably, in step 7, according to the texture feature and the semantic correspondence of the same - scale class semantic segmentation of the UV Map U seg , use the semantic mask to construct the semantic texture feature
[0022] Preferably, in step 8, use the semantic texture feature to replace the mean and variance of the semantic mask part of the semantic texture feature with the mean and variance of the semantic mask part, and input it into the pre - trained Decoder network for traditional texture style transfer to generate the semantic texture style transfer result, which together with the 3D shape model generated in step 4 constitutes an innovative 3D model of shape and texture.
[0023] In a second aspect, the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method described in the first aspect are implemented.
[0024] In a third aspect, the present invention provides a computer - readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the method described in the first aspect are implemented.
[0025] In a fourth aspect, the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.
[0026] The present invention uses a double - residual gated network and an Mlp network, with the mask, texture perception, and key - point correspondence relationships in the two - dimensional image as supervision signals, to learn the shape features of the source object and the target object and fuse them to generate a novel 3D model. And a semantic mask is introduced in the traditional texture style transfer to achieve partial style transfer and generate a richer texture map.
[0027] Compared with the prior art, the significant advantages of the present invention are as follows: (1) By using a dual-residual gated network and an Mlp network, and taking the mask, texture perception, and key-point correspondence relationship in a two-dimensional picture as supervision signals, the present invention learns the shape features of the source object and the target object and fuses them to generate a novel 3D model, solving the problem of realizing three-dimensional deformation starting from a single-view picture; (2) The present invention introduces a semantic mask in traditional texture style transfer. Through the style transfer of the texture features with the semantic mask added, partial style transfer is realized, and a richer texture map is generated; (3) Since cameras (such as mobile phones) are widely used in daily life, it is easier and cheaper to take single-view pictures than 3D data; therefore, instead of obtaining 3D input information, we use single-view pictures to solve the new task of generating 3D objects with novel shapes and textures; (4) The present invention proposes a three-dimensional deformation method based on conveniently obtained single-view pictures. This method takes two single-view pictures as input, extracts three-dimensional shape features, and realizes three-dimensional deformation; (5) The present invention proposes 3D shape and UV map texture transfer from the source image to the target image to create 3D objects. We design a shape transfer network to directly generate a novel 3D model and introduce a semantic UV map texture transmission method to obtain a new UV map texture. More importantly, the present invention can solve the task of generating a novel 3D model starting from a single-view picture.
[0028] The following further describes the present invention in detail with reference to the accompanying drawings. Description of the Drawings
[0029] Figure 1 It is a schematic flowchart of the present invention.
[0030] Figure 2 It is an overall network framework diagram of the present invention.
[0031] Figure 3 It is a framework diagram of the dual-residual gated network module in the present invention.
[0032] Figure 4 It is a framework diagram of the texture transfer module in the present invention.
[0033] Figure 5 It is a comparison diagram of the effects of the present invention in shape transfer and three-dimensional deformation transfer methods.
[0034] Figure 6 It is a comparison diagram of the effects of the texture transfer method with added semantics in the present invention and other texture transfer methods.
[0035] Figure 7 It is an effect diagram of texture transfer after adding different semantic partial control gates in the present invention.
[0036] Figure 8 This is the effect diagram of the present invention in simulating biological evolution.
[0037] Figure 9 This is the effect diagram of the present invention in controlling the change of the feature fusion ratio factor of the source object and the target object. Detailed implementation manners
[0038] As Figure 1 、 Figure 2 shown, a three-dimensional style transfer method algorithm based on a single-view image. Given the original image S and the target image T of the same category, use the Encoder of UMR to extract their shape features F S , F T , UV Map U S , U T and the camera pose P S , P T and the category semantic segmentation UV Map U seg ; use the Dual Residual Gated Network (DRG Net) to extract the shape features F S , F T in the shape features that can be fused with the same semantics Set the fusion ratio factor α to control the shape features with the same semantics The degree of feature fusion, and fuse the shape features according to the ratio Generate shape fusion features Use the Mlp network to take the shape fusion features as the input to obtain the three-dimensional space point coordinates that can combine the shape characteristics of the two images and add them to the three-dimensional category common template to generate a specific three-dimensional shape model; calculate the shape mask loss L mask , perceptual loss L per and three-dimensional key point correspondence loss L key , and optimize the parameters of the Dual Residual Gated Network and the Mlp network; use the pre-trained VGG network for traditional texture style transfer to extract the UV Map U S , U T texture features V S , V T ; downsample the category semantic segmentation UV Map U seg to the scale of the texture features V S , V T , and add the other semantic mask parts except the background part of U seg to the texture features V S , V T respectively to form each semantic texture feature Use the mean and variance of the semantic mask part of the semantic texture feature to replace the semantic texture feature The mean and variance of the semantic mask part are input into the trained Decoder network of traditional texture style transfer to generate the semantic texture style transfer result, which together with the above-generated 3D shape model constitutes an innovative 3D model of shape and texture. The specific steps are as follows:
[0039] Step 1: Given the original image S ∈ R 256×256×3 and the target image T ∈ R 256×256×3 of the same category, use the trained UMREncoder to obtain the camera pose C S ∈ R 7 , C T ∈ R 7 , the preliminary shape feature F S ∈ R 512 , F T ∈ R 512 and the UV Map texture map U S ∈ R 128×256×3 , U T ∈ R 128×256×3 , and the semantic texture UV Map segmentation map U seg ∈ {1, 2, 3, 4, 5} 128×256 . The 3D shape transfer network and the texture transfer network complete the generation of the novel 3D model based on the above data.
[0040] Step 2: As Figure 3 shown, the Dual Residual Gated Network (DRG Net) takes the preliminary shape feature F S ∈ R 512 , F T ∈ R 512 as the input and constructs a source branch and a target branch. Since these features are extracted from the same encoder, their coordinates have a potential correspondence. Therefore, the model designs a gated signal shared in the dual branches to select the features F S and F T in the same coordinates to facilitate shape transfer. Subsequently, the network is respectively connected to two single-layer perceptrons in the dual branches, and a residual connection is added to alleviate overfitting and gradient vanishing and prevent distortion caused by unselected features. The formula of the first dual residual gated unit for the initial input is described as follows:
[0041]
[0042]
[0043]
[0044] Among them, is the weight parameter of the l-th unit. The DRG Net iterates the double-residual gated unit L times to enhance the individual feature representation, progressively refine the features, and thus extract the shape features F S , F T and the shape features that can be fused with the same semantics in the middle
[0045] Step 3: Set the fusion ratio factor α to control the shape features with the same semantics and fuse the shape features according to the ratio Generate the shape fusion features according to the following formula
[0046]
[0047] Step 4: Use the Mlp network with the shape fusion features as the input for feature fusion. The specific structure is a simple two-layer neural network with a ReLU activation function. The specific formula is as follows:
[0048]
[0049] where the network parameter is the output of this layer. Based on this, the three-dimensional space point coordinates combining the shape characteristics of the two images can be obtained and added to the three-dimensional category common template to generate a specific three-dimensional shape model;
[0050] Step 5: Calculate the shape mask loss L mask , the perceptual loss L per and the three-dimensional key point correspondence loss L key , and optimize the parameters of the double-residual gated network and the Mlp network.
[0051] The shape mask loss L mask is to calculate the negative IoU between the real instance mask M and the predicted mask . The predicted mask is rendered from the generated three-dimensional model. L mask can be defined as:
[0052]
[0053]
[0054] where ⊙ represents element-wise multiplication.
[0055] The perceptual loss L per is to calculate the perceptual distance between the input image and the predicted image generated from its UV Map texture map. This loss can improve the visual quality of the texture by capturing details. The specific formula is as follows:
[0056]
[0057]
[0058] 3D Keypoint Correspondence Loss L key A 3D Keypoints Loss is proposed to achieve shape transfer between the source and the target. The 3D keypoints are obtained by detecting 2D keypoints from the image, determining their corresponding vertices on the reconstructed 3D shape using the predicted UV Map texture, and projecting them onto the symmetry plane. Since the symmetry plane is the same for all output 3D models, the distances between the keypoints are calculated based on their projections and are computed by the following formula:
[0059]
[0060] where N is the number of keypoints and λ is the balancing parameter. The shape transfer loss is generally summarized as follows:
[0061] L shape = L mask + L per + L key
[0062] Step 6: Use the trained VGG network for traditional texture style transfer to extract the UV Map U S , U T texture feature V S ∈ R 512×64×128 , V T ∈ R 512×64×128 .
[0063] Step 7: As shown in Figure 4 , downsample the class semantic segmentation UV Map U seg to the texture feature V S , V T scale, and add the other semantic mask parts of U seg except the background part to the texture feature V S , V T respectively to form each semantic texture feature Semantic style transfer integrates the semantic UV Map texture segmentation map mask U seg with any style transfer method for UV Map texture transfer. This method receives the source and target features V S , V T , and aligns the channel-wise means and variances of the masked V S without any additional parameters to match the masked V TThe mean and variance. This model extends AdaIN and linear style transfer (LST) as well as Exact Histogram Matching (EHM).
[0064] There are a total of 5 semantic parts in semantic segmentation. AdaIN is extended to semantic AdaIN (SAdaIN), and the corresponding AdaIN feature matrix is reconstructed according to the following formula:
[0065]
[0066]
[0067] where is a binary matrix that repeats the index matrix 512 times, with a dimension of 512×64×128. ⊙ is element-wise multiplication, σ is the variance, and μ is the mean. The with index 5 represents the non-semantic part and does not contain semantic information.
[0068] Similar to SAdaIN, LST is extended to semantic LST (SLST) with the following formula:
[0069]
[0070]
[0071] Similarly, semantic EFDM (SEFDM) is defined as follows:
[0072]
[0073]
[0074] Table 1 Quantitative comparison between 3D geometric deformation methods
[0075] Method NC DSN KPD NT Ours Mask IoU↑ 0.6670 0.6937 0.5699 0.5209 0.7316
[0076] Table 1 is a quantitative comparison of the results of shape transfer in the present invention with other 3D geometric deformation methods. The mask IoU metric is used to quantify the quality of the shape transformation. In terms of preserving the shape contours of the source and target, our results are better than the comparison methods. Note that a mask IoU in the range of 0.7 to 0.95 may be reasonable, because a smaller IoU indicates that the result does not inherit the shape information of the target and source birds, while a higher IoU indicates that the result inherits the shape information, that is, there is no evolutionary diversity. In addition, our texture transformation aims to improve the texture diversity of the species, and we have not found a suitable metric to measure this.
[0077] Table 2 User survey
[0078]
[0079] To better evaluate the performance of our model compared with existing models, a user study was conducted. It consists of three main parts: shape transfer comparison, texture transfer comparison, and authenticity judgment, with a total of 25 questions. As a result, we collected 102 questionnaire responses, totaling 2,550 votes. Table 2 shows that in the shape transfer comparison, 52.5% of users preferred our results, compared with 27.8% for NC, 11.8% for DSN, 3.5% for KPD, and 4.3% for NT. In the texture transfer comparison, 76.9% of users preferred the results of SAdaIN over AdaIN, 71.2% preferred the results of SLST over LST, and 64.1% preferred the results of SEFDM over EFDM. In the authenticity judgment, 73.7% of users considered our results to be relatively realistic.
[0080] Appendix Figure 5 Shows the results of shape transfer. We can observe that our method achieves reasonable shape transfer and can better match the shape features of the source object and the target object. It shows that when the shape difference between the source object and the target object of different species is large, the comparison method will get some unreasonable distortions. A possible reason is that it is difficult to align the semantic parts. In contrast, our method learns to reconstruct the 3D model to prevent shape deformation and at the same time unfolds shape transfer to generate new shapes.
[0081] Appendix Figures 6 - 7 Shows the results of the style transfer algorithm (such as AdaIN, LST, and EFDM) after using the semantic texture transfer module. Figure 7 Shows the results after adding different semantic part control gates. We can see that semantic style transfer improves the influence of style transfer on each semantic part of all algorithms because the semantic mask prevents the influence between different semantic parts. In addition, in Figure 7 , the semantic part control gate further increases the diversity of the results according to the semantic mask, making our semantic texture transfer more in line with the law of natural evolution.
[0082] Appendix Figure 8 Shows that to further verify the effectiveness of our model, we tried to automatically generate the morphological evolution of birds between the same species for biologists to study. We collected real hybrid birds and their parents from the Internet. Our model uses the parental species to simulate hybridization as the source object and the target object input, and the synthetic results are very similar to the real hybrid examples. It should be noted that although our results are far from the goals of biologists, this is a meaningful attempt to extend style transfer to animal morphological evolution.
[0083] Appendix Figure 9 It shows that using α adjusts the fusion ratio between the source and target features from -1 to 1. It indicates that the scale parameter α effectively controls the presentation of the source and target features in the result. When α = -1 or 1, the result is exactly the source or target 3D reconstruction. When α = 0, the result is a combination of half of the shape features of the source object and half of the shape features of the target object.
[0084] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A three-dimensional style transfer method for single-view images, characterized in that, generate a three-dimensional model that fuses the shape and texture characteristics of the original image and the target image of the same category, including the following steps: Step 1: Given the original image S and the target image T of the same category, use the Encoder of UMR to extract their shape features F S , F T , UV Map U S , U T , as well as the camera pose P S , P T , and the category semantic segmentation UV Map U seg ; Step 2. Use the double residual gated network to extract the shape features F S , F T and the shape features that can be fused with the same semantics Step 3: Set a fusion ratio factor α to control the shape features with the same semantics Degree of feature fusion, and fuse the shape features according to the ratio Generate shape fusion features Step 4: Use the Mlp network with the shape fusion feature as the input to obtain the three-dimensional spatial point coordinates that can combine the shape characteristics of the two images and add them to the three-dimensional category common template to generate a specific three-dimensional shape model; Step 5: Calculate the shape mask loss $L$ mask , the perceptual loss $L$ per and the 3D key point correspondence loss $L$ key , and optimize the parameters of the double residual gating network and the Mlp network; Step 6: Extract the UV Map u using the pre-trained VGG network for traditional texture style transfer S , u T texture feature V S , V T ; Step 7: Downsample the category semantic segmentation UV Map U seg to the texture feature V S , V T at a scale, and separately add the other semantic mask parts except the background part to the texture feature V seg , V S to form respective semantic texture features T Step 8: Use semantic texture features Replace the semantic texture features with the mean and variance of the semantic mask part The mean and variance of the semantic mask part, and input them into the trained Decoder network of traditional texture style transfer to generate the semantic texture style transfer result, which together with the three-dimensional shape model generated in Step 4 constitutes an innovative three-dimensional model of shape and texture.
2. The three-dimensional style transfer method for single-view images according to claim 1, characterized in that, In step 1, the Encoder of UMR is used to extract the shape features F of the original image S and the target image T respectively S ,F T As the preliminary shape feature, the UVMap U S ,U T and the category semantic segmentation UV Map U seg have a fixed mapping method with the 3D model space points and textures and are semantically consistent 3. The three-dimensional style transfer method based on single-view images according to claim 1, characterized in that, In the double-residual gated network in step 2, the shape features of the original image and the target image are hierarchically extracted using an iterative structure, the original information of the shape features is maintained using a residual structure, and gated information is generated based on the two shape features to obtain the residual change of the shape features.
4. The three-dimensional style transfer method based on single-view images according to claim 1, characterized in that, The scaling factor in step 3 controls the fusion ratio of the shape features of the original image and the target image.
5. The three-dimensional style transfer method based on single-view images according to claim 4, characterized in that, Generate shape fusion features according to the following formula 6. The three-dimensional style transfer method based on single-view images according to claim 1, characterized in that, In step 5, the generated innovative 3D model is respectively projected onto the 2D plane in the original image camera pose and the target image camera pose, and the IOU is calculated between the mask after projection and the masks of the original image and the target image to generate L mas , and the perceptual loss L is calculated between the textured image after projection onto the 2D plane and the original image and the target image per , and the key point loss L is calculated for the key points after projection key ; Through L mask and L per , the overall shape of the 3D model is controlled, and L key can maintain the relative positions between the various parts of the 3D model.
7. The three-dimensional style transfer method based on single-view images according to claim 1, characterized in that, In step 7, segment the UV Map U according to the texture features and the class semantic segmentation at the same scale seg Semantic correspondence, use the semantic mask to construct the semantic texture features 8. An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the steps of the method described in any one of claims 1-7 are implemented.
9. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the program is executed by a processor, the steps of the method described in any one of claims 1-7 are implemented.
10. A computer program product, including a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method described in any one of claims 1-7 are implemented.
Citation Information
Patent Citations
Image processing method and device, storage medium and electronic equipment
CN113780326A
Image style migration system and method based on three-branch clustering semantic segmentation
CN113902613A