Three-dimensional model stylization method and apparatus, electronic device, and storage medium
Patent Information
- Application Number
- CN202111074530.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-14
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2041-09-14
AI Technical Summary
然而,如果要将一张目标图像的风格迁移到三维模型上,由于三维模型是三维的,而目标图像是二维的,无法使用三维卷积进行处理
[0022] In a sixth aspect, embodiments of this disclosure also provide a computer-readable medium having a computer program stored thereon, characterized in that, when executed by a processor, the program implements the three-dimensional model stylization method as described in the first or second aspect.
Smart Images

Figure CN115810101B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image processing technology, and more particularly to a three-dimensional model stylization method, apparatus, electronic device, and storage medium. Background Technology
[0002] Stylization, also known as style transfer, transfers the style of an artistic image to a regular 2D image, giving the 2D image a unique artistic style while retaining its original content, such as cartoon, comic, oil painting, watercolor, or ink painting. Currently, deep learning networks can be used for stylizing 2D images. However, transferring the style of a target image to a 3D model is problematic because the 3D model is three-dimensional while the target image is two-dimensional; 3D convolution cannot be used for this purpose. There is currently no effective solution for stylizing 3D models. Summary of the Invention
[0003] This disclosure provides a method, apparatus, electronic device, and storage medium for stylizing three-dimensional models to achieve stylization of three-dimensional models.
[0004] In a first aspect, embodiments of this disclosure provide a three-dimensional model stylization method, including:
[0005] Obtain the 3D model to be stylized and the target image for stylization;
[0006] The three-dimensional model is rendered using a preset network to obtain a two-dimensional rendered image and spatial feature parameters of pixels. Based on the spatial feature parameters and the stylized target image, the texture features in the two-dimensional rendered image are stylized to obtain a stylized three-dimensional model.
[0007] Secondly, embodiments of this disclosure also provide a method for stylizing a three-dimensional model, including:
[0008] Scan at least two two-dimensional input images, each of which includes features of the target to be modeled from a corresponding viewpoint;
[0009] A three-dimensional model of the target to be modeled is established based on each of the two-dimensional input images;
[0010] The three-dimensional model is stylized based on the preset network, the stylized target image, and the spatial feature parameters of the pixels in the three-dimensional model.
[0011] Thirdly, embodiments of this disclosure also provide a three-dimensional model stylization apparatus, comprising:
[0012] The acquisition module is used to acquire the 3D model to be stylized and the target image for stylization.
[0013] The stylization module is used to render the 3D model through a preset network to obtain a 2D rendered image and spatial feature parameters of pixels, and to stylize the texture features in the 2D rendered image according to the spatial feature parameters and the stylization target image to obtain a stylized 3D model.
[0014] Fourthly, embodiments of this disclosure also provide a three-dimensional model stylization apparatus, comprising:
[0015] The scanning module is used to scan at least two two-dimensional input images, each of which includes features of the target to be modeled from a corresponding viewpoint;
[0016] The modeling module is used to establish a three-dimensional model of the target to be modeled based on each of the two-dimensional input images;
[0017] The execution module is used to stylize the 3D model according to the preset network, the stylized target image, and the spatial feature parameters of the pixels in the 3D model.
[0018] Fifthly, embodiments of this disclosure also provide an electronic device, including:
[0019] One or more processors;
[0020] Storage device for storing one or more programs;
[0021] When the one or more programs are executed by the one or more processors, the one or more processors implement the three-dimensional model stylization method as described in the first or second aspect.
[0022] In a sixth aspect, embodiments of this disclosure also provide a computer-readable medium having a computer program stored thereon, characterized in that, when executed by a processor, the program implements the three-dimensional model stylization method as described in the first or second aspect.
[0023] This disclosure provides a method, apparatus, electronic device, and storage medium for stylizing a 3D model. The method includes: acquiring a 3D model to be stylized and a target image for stylization; rendering the 3D model using a preset network to obtain a 2D rendered image and spatial feature parameters of pixels; and stylizing the texture features in the 2D rendered image based on the spatial feature parameters and the target image for stylization to obtain a stylized 3D model. This technical solution renders the 3D model as a 2D image and considers the spatial features of each pixel. All pixels in the 2D rendered image, including adjacent pixels with discontinuous texture features, can be stylized, ensuring the consistency of the 3D model's spatial structure before and after stylization. Attached Figure Description
[0024] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0025] Figure 1 This is a flowchart of the three-dimensional model stylization method in Embodiment 1 of this disclosure;
[0026] Figure 2 This is a flowchart of the three-dimensional model stylization method in Embodiment 2 of this disclosure;
[0027] Figure 3 This is a schematic diagram of the sphere model viewed from multiple perspectives in Embodiment 2 of this disclosure;
[0028] Figure 4 This is a schematic diagram of the three-dimensional model stylization process in Embodiment 2 of this disclosure;
[0029] Figure 5 This is a flowchart of the three-dimensional model stylization method in Embodiment 3 of this disclosure;
[0030] Figure 6 This is a schematic diagram of the structure of the three-dimensional model stylization device in Embodiment 4 of this disclosure;
[0031] Figure 7 This is a schematic diagram of the structure of the three-dimensional model stylization device in Embodiment 5 of this disclosure;
[0032] Figure 8 This is a schematic diagram of the hardware structure of the electronic device in Embodiment 5 of this disclosure. Detailed Implementation
[0033] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0034] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0035] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0036] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0037] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0038] In the following embodiments, each embodiment provides optional features and examples. The various features described in the embodiments can be combined to form multiple optional solutions. Each numbered embodiment should not be regarded as only one technical solution. Furthermore, unless otherwise specified, the embodiments and features in the embodiments of this disclosure can be combined with each other.
[0039] Example 1
[0040] Figure 1 This is a flowchart of a 3D model stylization method according to Embodiment 1 of this disclosure. This method is applicable to the stylization of 3D models. Specifically, the 3D model is input into a preset network, which then performs comprehensive stylization of the 3D model based on the style of the target image, thereby outputting a stylized 3D model with the same structure as the original 3D model. This method can be executed by a 3D model stylization device, which can be implemented in software and / or hardware and integrated into an electronic device. In this embodiment, the electronic device can be a computer, laptop, server, tablet computer, or smartphone, or other device with image processing capabilities.
[0041] It should be noted that the process of stylizing a 3D model can be understood as stylizing the texture of the 3D model's surface. This process requires rendering the 3D model as a 2D image containing texture features. The following issues need to be addressed: When rendering the portion visible from the 3D model's surface from a certain viewpoint into a 2D image, a projection plane needs to be determined. However, some locations within the visible portion have excessively large angles with the projection plane. These locations, after being projected onto the projection plane and stylized, will exhibit significant deformation relative to the original 3D model. Because the 3D model is three-dimensional, it's impossible to render all the textures of the 3D model onto the 2D image at once. If multiple renderings are performed, the stylization effect of the texture features in each rendered 2D image will have poor continuity. Furthermore, because the 3D model is three-dimensional and complex in shape, occlusion may occur at different depths from a certain viewpoint. Therefore, adjacent pixels projected onto the 2D image may not be adjacent in their actual positions on the 3D model, and their texture features may not be continuous, making stylization difficult. For example, when looking at a person's head, you can see the lowest point of the chin and then the point of the neck. From the front, these two points are adjacent, but in fact, the two points are at different depths and their texture features are not continuous. When stylizing, the difference in texture features between these two points needs to be considered so that the different depths of the two points can still be reflected after stylization.
[0042] like Figure 1 As shown in Embodiment 1 of this disclosure, a three-dimensional model stylization method specifically includes the following steps:
[0043] S110. Obtain the 3D model to be stylized and the target image for stylization.
[0044] In this embodiment, the 3D model to be stylized can be any solid model, such as a 3D geometric model, a model generated from an entity (e.g., a model built from a table), a model built from multiple 2D images (e.g., photos of a table taken from different angles), a virtual model (e.g., a table model synthesized using software), etc. The object it represents can be a real-world entity or a fictional object.
[0045] The stylized target image is two-dimensional and can be understood as a reference image with a specific artistic style. The purpose of 3D model stylization is mainly to transfer the specific artistic style from the stylized target image to the surface of the 3D model, so that the surface of the 3D model has corresponding texture patterns, color patterns, or visual perceptions. The stylized target image can be downloaded from an online image library or input or specified by the user.
[0046] S120. The three-dimensional model is rendered through a preset network to obtain a two-dimensional rendered image and spatial feature parameters of pixels. Based on the spatial feature parameters and the stylized target image, the texture features in the two-dimensional rendered image are stylized to obtain a stylized three-dimensional model.
[0047] In this embodiment, the preset network is primarily a deep neural network with the following two functions: rendering a 3D model into a 2D rendered image containing texture features, and stylizing the texture features in the 2D rendered image and reflecting the stylized texture features at corresponding positions on the surface of the 3D model. The input to the preset network is the 3D model to be stylized and the target image for stylization. The 3D model provides content features, and the target image provides style features. Combining the content features and style features yields the stylized 3D model as the output.
[0048] Spatial feature parameters can include the angle between the normal and the viewpoint for each pixel, the pixel depth, etc., which are used to determine the correspondence between pixels in the 2D rendered image and pixels on the surface of the 3D model, thereby stylizing the texture features in the 2D rendered image onto the surface of the 3D model.
[0049] Specifically, the rendering process can be implemented using a renderer, while the stylization process can be implemented using a stylization network. The pre-defined network first renders the 3D model using the renderer, drawing it onto a projection plane to obtain a 2D rendered image. During this process, the renderer can also calculate the spatial feature parameters of the pixels. Then, the 2D rendered image passes through the stylization network. In the stylization network, the correspondence between pixels in the 2D rendered image and pixels on the surface of the 3D model can be determined based on the spatial feature parameters. Therefore, by using the style of the target image, the texture features of the corresponding pixels on the surface of the 3D model can be stylized, i.e., style transfer. Based on this, the stylization of the entire 3D model can be gradually achieved through multiple renderings from different perspectives.
[0050] Specifically, the stylization process can be implemented using a neural network with an encoder-decoder structure. The encoder receives the input 2D rendered image and the stylized target image, and extracts the feature vectors of both. These feature vectors can be understood as another representation of the features and information of the input. The decoder outputs an expected result based on these feature vectors. This expected result is the 3D model obtained by transferring the style of the stylized target image to the surface of the 3D model corresponding to the 2D rendered image.
[0051] Optionally, the renderer can be a differentiable renderer, which, after training, can learn the rules for obtaining a two-dimensional rendered image from a three-dimensional model.
[0052] Stylization networks can be image segmentation networks, such as U-Net. U-Net uses a network structure that includes downsampling and upsampling. The main purpose of downsampling is to gradually present the texture patterns of each pixel and its surrounding pixels based on the image features (which can also be understood as style features) of the stylized target image. The main purpose of upsampling is to combine the downsampling information with the features of the original 3D model (i.e., the 3D model to be stylized) to restore the details of the 3D model and gradually restore the resolution to the level of the original 3D model.
[0053] This embodiment provides a 3D model stylization method that renders the 3D model into a 2D image and utilizes the spatial feature parameters of pixels to clarify the spatial characteristics and spatial positional relationships of each pixel in the 3D model. This solves the three problems mentioned earlier: reducing deformation during stylization, considering the impact of discontinuous texture features of adjacent pixels on stylization, and ensuring the effectiveness of 3D model stylization in all directions. Furthermore, while stylizing all pixels in the 2D rendered image, including adjacent pixels with discontinuous texture features, the method ensures the consistency of the spatial structure of the 3D model before and after stylization.
[0054] Example 2
[0055] Figure 2 This is a flowchart of the three-dimensional model stylization method in Embodiment 2 of this disclosure. Based on the above embodiments, Embodiment 2 specifies the process of rendering and stylizing a three-dimensional model using a preset network.
[0056] In this embodiment, a 3D model is rendered using a preset network to obtain a 2D rendered image and spatial feature parameters of pixels. Based on the spatial feature parameters and a stylization target image, the texture features in the 2D rendered image are stylized. This includes: determining multiple viewpoints that can cover the 3D model; rendering the 3D model from the current viewpoint using the preset network to obtain a 2D rendered image corresponding to the current viewpoint and spatial feature parameters of the corresponding pixels; stylizing the texture features in the 2D rendered image corresponding to the current viewpoint based on the spatial feature parameters of the corresponding pixels and the stylization target image; selecting the next viewpoint as the current viewpoint and repeating the rendering and stylization operations for the current viewpoint until a complete stylized 3D model is obtained. Based on this, the task of stylizing the 3D model from multiple viewpoints is decomposed into multiple sequential processes. The rendering and stylization of each viewpoint are performed on the basis of the rendering and stylization of the previous viewpoint, thereby gradually completing the overall stylization of the 3D model, ensuring the continuity of the stylized 3D model and the consistency of the spatial structure of the 3D model before and after stylization.
[0057] In this embodiment, the preset network includes a renderer and a stylization network. The spatial feature parameters of the corresponding pixel in the current viewpoint include the angle between the normal of the corresponding pixel in the current viewpoint and the viewpoint, as well as the depth of the corresponding pixel in the current viewpoint. If the current viewpoint is not the first viewpoint, the spatial feature parameters of the corresponding pixel in the current viewpoint also include the mask of the part of the 2D rendered image corresponding to the current viewpoint that has been stylized in the previous viewpoint. That is, for non-first viewpoints, the input of the renderer is the partially stylized 3D model output in the previous viewpoint. Based on the spatial feature parameters of the corresponding pixel in the current viewpoint, the spatial positional relationship of each pixel in the 3D model can be clearly defined, thereby achieving effective stylization of all pixels in the 2D rendered image, including adjacent pixels with discontinuous texture features.
[0058] It should be noted that if the world coordinate system is used to locate the pixels on the surface of the 3D model, the projection surface is fixed under different viewpoints, so the normal corresponding to each pixel remains unchanged. However, during rendering and stylization, the line of sight to each pixel from the human eye or camera is different. If the camera coordinate system is used to locate the pixels on the surface of the 3D model, the line of sight to any pixel can be used as a reference. That is, when looking at a certain pixel, that pixel can be moved to the point of gaze of the line of sight. In this case, the 3D model and its projection surface are usually rotated, so the corresponding normal is different when looking at different pixels. Regardless of which coordinate system is used, when rendering and stylizing the 3D model from the current viewpoint, the angle between the normal corresponding to each pixel and the line of sight is determined. This embodiment uses the world coordinate system as an example for explanation.
[0059] like Figure 2 As shown in Embodiment 2 of this disclosure, a three-dimensional model stylization method includes the following steps:
[0060] S210. Obtain the 3D model to be stylized and the target image for stylization.
[0061] S220. Determine the field of view to cover multiple perspectives of the 3D model.
[0062] Specifically, the viewpoint can be understood as the angle between the line of sight looking at a certain pixel on a projection surface and the perpendicular direction of the projection surface; the perpendicular direction of the projection surface is the normal; the field of view can be understood as the range that the eye or camera can capture when looking at the 3D model from a certain viewpoint. In order to ensure the stylization of the 3D model in all aspects, the field of view of multiple viewpoints needs to be able to cover all positions on the surface of the 3D model.
[0063] Figure 3 This is a schematic diagram of the sphere model viewed from multiple perspectives in Embodiment 2 of this disclosure. For example... Figure 3As shown, the 3D model is a sphere. From the front view, half of the sphere's surface is visible. This half is divided into left and right parts, each representing a quarter of the sphere's surface. These two parts are denoted as A and B. The view from the right side includes B, as well as a portion not visible from the front view, which is also a quarter of the sphere's surface and is denoted as C. The view from the rear side includes C, as well as a portion not visible from either the front or right view, which is also a quarter of the sphere's surface and is denoted as D. The view from the left side includes D and A. The field of view of these four perspectives covers all areas of the 3D model's surface.
[0064] S230. Render the 3D model from the current viewpoint using the renderer to obtain the 2D rendered image corresponding to the current viewpoint and the spatial feature parameters of the corresponding pixels.
[0065] Specifically, for the current viewpoint, the renderer renders the 3D model to obtain the corresponding 2D rendered image. This 2D rendered image contains the texture features of the 3D model surface that can be seen from the current viewpoint. The renderer can also calculate the spatial feature parameters of each pixel in the 2D rendered image. The spatial feature parameters of each pixel can include the angle between the line of sight from the current viewpoint to the pixel and the normal under the current viewpoint, as well as the depth of each pixel under the current viewpoint. The depth can be understood as the distance between the pixel and the eye or camera, or the distance between the pixel and the projection surface. The depth can reflect the position of the pixel in the line of sight direction.
[0066] It should be noted that for each viewpoint other than the first viewpoint, the spatial feature parameters of the corresponding pixel also include the mask of the part that has been stylized in the previous viewpoint in the 2D rendered image corresponding to that viewpoint. This mask is used to distinguish between the stylized and unstylized parts, so that the stylization network can stylize the unstylized parts.
[0067] S240. Using a stylization network, the texture features in the two-dimensional rendered image corresponding to the current viewpoint are stylized based on the spatial feature parameters of the corresponding pixels at the current viewpoint and the stylization target image.
[0068] Optionally, the field of view of adjacent viewpoints overlaps; stylizing the texture features in the 2D rendered image corresponding to the current viewpoint includes: stylizing the texture features of the parts of the 2D rendered image corresponding to the current viewpoint that did not appear in the 2D rendered image corresponding to the previous viewpoint.
[0069] refer to Figure 3The sphere model shown has overlapping fields of view between the front and right sides, and between the right and rear sides, but the left side view is not considered. Furthermore, at the boundary between D and A, to ensure stylistic continuity using spatial features, the fields of view of the rear side can overlap with those of the front side, meaning D and A overlap. This can be achieved by rotating the rear side view counter-clockwise a certain angle towards the front side view. Based on this, if the current viewpoint is a frontal view, A and B are rendered to obtain a 2D rendered image, and their texture features are stylized. Then, if the current viewpoint is a right-side view, B and C are rendered to obtain a 2D rendered image, and their texture features are stylized. Since B has already been rendered and stylized in the frontal view, considering the spatial relationship between B and C, only C needs to be stylized. Next, if the current viewpoint is a rear-side view, C and D (including the portion overlapping with A) are rendered to obtain a 2D rendered image, and their texture features are stylized. Since C has already been rendered and stylized in the right-side view, and the portion of D overlapping with A has already been stylized in the frontal view, only the portion of D excluding A needs to be stylized. This avoids repeatedly stylizing the same part, improving the efficiency of stylization.
[0070] S250: Is the entire 3D model stylized? If yes, then execute S270; otherwise, execute S260.
[0071] In this embodiment, determining whether the entire 3D model has been stylized can be replaced by determining whether the current viewpoint is the last viewpoint. If so, it means that the 3D model stylization is complete, and in this case, the output of the stylization network is the stylized model; if not, it is necessary to select the next viewpoint and continue to perform rendering and stylization operations.
[0072] S260. Select the next viewpoint as the current viewpoint and return to S230.
[0073] S270, outputs a complete stylized 3D model.
[0074] Figure 4 This is a schematic diagram of the three-dimensional model stylization process in Embodiment 2 of this disclosure. For example... Figure 4As shown, the preset network 21 includes a renderer 211 and a stylization network 212. The 3D model 22 to be stylized is input to the renderer 211, and the stylization target image 23 is input to the stylization network 212. The output of the stylization network 212 is a stylized 3D model 24. The renderer 211 renders the 3D model 22 from the selected current viewpoint to obtain a corresponding 2D rendered image, and calculates the spatial feature parameters of the corresponding pixels. The spatial feature parameters include the angle between the normal of each pixel and the viewpoint, and the depth of each pixel. If the current viewpoint is not the first viewpoint, the spatial feature parameters also include the mask of the part that has been stylized in the previous viewpoint. The spatial feature parameters of the current viewpoint and the 2D rendered image are combined and input to the stylization network 212. The stylization network 212 is used to determine the relationship between the pixels in the 2D rendered image and the pixels on the surface of the 3D model based on the spatial feature parameters of the current viewpoint, thereby combining the image features of the stylization target image 23 to stylize the texture features of the corresponding part of the 3D model under the current viewpoint.
[0075] If the current viewpoint is the last viewpoint, the stylization network 212 outputs a stylized 3D model 24; if the current viewpoint is not the last viewpoint, the stylization network 212 outputs a partially stylized 3D model, which is then re-inputted to the renderer 211 from a newly selected viewpoint. It is evident that rendering and stylizing different viewpoints is a sequential process. That is, for each viewpoint, only a corresponding portion of the 3D model can be stylized after each rendering. The rendering and stylization operations in each viewpoint are based on the rendering and stylization of previous viewpoints, continuing until the last viewpoint, where a final rendering and stylization is performed, and the output is denoted as the stylized 3D model.
[0076] The following examples illustrate the 3D model stylization process:
[0077] Assuming there are n viewpoints (n≥2, n is a positive number), the first viewpoint (n=1) is selected and denoted as V1. The 3D model is rendered from V1 using a renderer to obtain the corresponding 2D rendered image I1. I1 contains the texture features F1 of the 3D model surface that can be seen from V1. The renderer can also obtain the spatial feature parameters S1 of the pixels in I1. S1 can include the angle α between the line of sight from V1 to each pixel i and the normal of V1. i1 And the depth D of each pixel i in the current view. i1 The stylization network stylizes the texture features of the 3D model surface visible from the current viewpoint based on the stylization target image I0, F1, S1, and I1, and outputs R1;
[0078] Select the next viewpoint V2 (n=2), and render the 3D model from V2 using the renderer to obtain the corresponding 2D rendered image I2. I2 contains the texture features F2 of the 3D model surface that can be seen from V2. The renderer can also obtain the spatial feature parameters S2 of the pixels in I2. S2 can include the angle α between the line of sight from V2 to each pixel i and the normal of V2. i2 The depth D of each pixel i in the current view. i2 The stylization network uses I0, F2, S2, M1, and I2 to stylize the texture features of the 3D model surface visible from the current viewpoint, and outputs R2. R2 includes not only the parts that have been stylized in viewpoint V1, but also the parts that are stylized in viewpoint V2.
[0079] If V2 is the last viewpoint, then R2 is the stylized result; otherwise, continue to select the next viewpoint V3 and repeat the above rendering and stylization operations until the complete stylized 3D model output from the last viewpoint is obtained.
[0080] It should be noted that for each perspective other than the first perspective, V k (k≥2, n is a positive number), the spatial feature parameters of the corresponding pixel also include those in V. k Corresponding 2D rendered image I k In the middle, from the perspective V k-1 Mask M of the stylized portion k-1 .
[0081] In addition, in order to ensure the effect of stylizing the texture features of pixels at the intersection of different viewpoints, the field of view of adjacent viewpoints can overlap, that is, usually three or more viewpoints are selected.
[0082] In one embodiment, the method further includes: training a preset network based on a sample model and a stylized target image until the value of the loss function of the preset network meets the requirements; wherein the value of the loss function is determined based on the values of the following three functions:
[0083] Content loss function is used to evaluate the loss between the stylization result and the sample model;
[0084] The style loss function is used to evaluate the loss between the stylization result and the stylized target image;
[0085] The continuity loss function is used to evaluate the super-resolution test sequence (Visual Geometry Group, VGG) loss between the stylized result and the sample model whose previous viewpoint was stylized.
[0086] In this embodiment, "content" mainly refers to the structure and outline of the sample model or stylization result, and the Euclidean distance can be used as an indicator to measure the content difference between the sample model and the stylization result. "Style" mainly refers to the texture patterns, color patterns, or visual perception of the sample model or stylization result, and the style difference between the sample model and the stylization result can be represented by the Gram matrix between the feature maps of the same hidden layer. "Continuity" mainly refers to the continuity of texture features between the stylization result and the sample model whose previous viewpoint was stylized, and can be represented using VGG loss.
[0087] The stylization result can be understood as a stylized sample model. A pre-trained network can be trained using the sample model and the stylized target image, allowing it to learn the rules governing the stylization result obtained from the sample model and the stylized target image. This allows for practical application in the stylization of 3D models. The loss function used during training can be one of the three loss functions mentioned above, such as the sum of the three loss functions or a weighted sum, ensuring that the stylized result is similar in content to the sample model, similar in style to the stylized target image, and that the VGG loss between the output stylized result and the previous viewpoint (the penultimate viewpoint) is minimized.
[0088] The training process is as follows: The sample model is rendered and stylized using an initial preset network to obtain the stylized result for the current viewpoint; the content loss L between this stylized result and the sample model is then calculated. Content The style loss L between the stylization result and the stylization target image Style The VGG loss L between this stylization result and the stylization result from the previous perspective. VGG By continuously training and adjusting the network parameters in the preset network, the overall loss is minimized, thereby optimizing the performance of the preset network and achieving good robustness. The overall loss function is, for example, L = L... Content +L Style +L VGG .
[0089] In one embodiment, the style loss function is a weighted sum of the losses between the image features of each pixel in the stylization result and the image features of the stylized target image, wherein the weight corresponding to each pixel is negatively correlated with the angle between the normal corresponding to each pixel and the line of sight.
[0090] In this embodiment, when calculating the style loss function value between the stylization result and the sample model, the loss between the image features of each pixel and the image features of the stylization target image is weighted. Specifically, the weight of each pixel is related to the angle between the normal and the viewpoint of that pixel. For example, the larger the angle between the normal and the viewpoint of that pixel, the more the pixel is off from the current viewpoint, and the smaller the corresponding weight should be. Thus, for the parts with large deformation when the sample model is drawn into a two-dimensional rendering image, its influence on style loss can be reduced.
[0091] In one embodiment, the image features of each pixel in the stylization result are determined by performing weighted convolution on the surrounding pixels of each pixel, wherein the weight of the convolution is negatively correlated with the depth difference between the pixel and the surrounding pixels.
[0092] In this embodiment, during the calculation of the style loss function, weighted convolution is used to extract features from the 2D rendered image or the stylized target image. For a pixel, the convolution weight is related to the depth difference between that pixel and its surrounding pixels; the greater the depth difference, the smaller the corresponding weight should be. Based on this, it can be ensured that discontinuous positions in the sample model remain discontinuous after stylization.
[0093] In one embodiment, the continuity loss function is the weighted sum of the VGG loss between each pixel in the stylization result and the corresponding pixel in the sample model of the previous viewpoint that was stylized, wherein the weight corresponding to each pixel is positively correlated with the angle between the normal corresponding to each pixel and the line of sight.
[0094] In this embodiment, when calculating the VGG loss function value between the stylized result and the sample model stylized in the previous viewpoint, the VGG loss of each pixel is assigned a weight. Specifically, the weight of each pixel is related to the angle between the normal and the line of sight to that pixel. For example, the larger the angle between the normal and the line of sight to that pixel, the more the pixel is away from the current viewpoint, and the greater the corresponding weight should be. Thus, the continuity of the parts with large deformation when the sample model is drawn into a two-dimensional rendering image can be given priority to reduce deformation problems.
[0095] This embodiment provides a 3D model stylization method. Before inputting a 2D rendered image into the stylization network, it performs differentiable rendering and calculation of spatial feature parameters to clarify the spatial positional relationship of each pixel in the 2D rendered image within the 3D model, ensuring the consistency of the 3D model's spatial structure before and after stylization. By decomposing the 3D model stylization task into multiple sequential processes from multiple perspectives, the rendering and stylization of each perspective are based on the rendering and stylization of the previous perspective, thus gradually completing the stylization of the overall 3D model, ensuring the continuity of the stylized 3D model and the consistency of the 3D model's spatial structure before and after stylization. By comprehensively evaluating the loss between the stylization result and the sample model based on the style loss function, content loss function, and continuity loss function, and assigning weights to the style loss, convolution, and VGG loss corresponding to different pixels, it reduces deformation during the stylization process, considers the impact of discontinuous texture features of adjacent pixels on stylization, and achieves high-quality stylization of the 3D model in all directions.
[0096] Example 3
[0097] Figure 5 This is a flowchart of the 3D model stylization method in Embodiment 3 of this disclosure. This method is applicable to situations where a 3D model is created and stylized based on multiple 2D input images. The electronic device in this embodiment can be a computer, laptop, server, tablet computer, or smartphone, or other device with image processing capabilities. For details not covered in this embodiment, please refer to the above embodiments.
[0098] like Figure 5 As shown in Embodiment 3 of this disclosure, a three-dimensional model stylization method includes the following steps:
[0099] S310. Scan at least two two-dimensional input images, each of which includes features of the target to be modeled from the corresponding viewpoint.
[0100] In this embodiment, at least two 2D input images are used to reflect the shape, color, texture, and other features of the same target to be modeled from different viewpoints, providing a basis for building a 3D model. To ensure accurate modeling, at least two 2D input images need to contain feature information of all locations on the surface of the target to be modeled. The 2D input images can be downloaded from an online image library or input or specified by the user. For example, if the target to be modeled is a table, photos are taken at the same horizontal height in a counterclockwise direction, taking one photo every 60 degrees, ensuring that the features of the same location of the target can be found in photos from adjacent viewpoints. The resulting multiple photos can then be used as 2D input images.
[0101] S320. Establish a three-dimensional model of the target to be modeled based on each of the two-dimensional input images.
[0102] In this embodiment, the process of building a three-dimensional model from a two-dimensional input image can also be understood as the three-dimensionalization of the two-dimensional input image. Based on multiple two-dimensional input images from different perspectives, the three-dimensional structure of the target to be modeled can be recovered. For example, for the multiple photos mentioned above, the three-dimensional model of the target to be modeled, i.e., the three-dimensional model to be stylized, is determined based on the shooting perspective, the two-dimensional coordinates of each pixel in the photo, and the relationship between the same pixel and its surrounding pixels in photos from different perspectives.
[0103] S330. Stylize the three-dimensional model according to the preset network, the stylized target image, and the spatial feature parameters of the pixels in the three-dimensional model.
[0104] In this embodiment, the preset network can be a pre-trained deep neural network. Its inputs are the 3D model to be stylized and the target image for stylization. The 3D model provides content features, while the target image provides style features. Combining the content features and style features yields a stylized 3D model as the output. During the stylization process, the preset network can utilize the spatial feature vectors of pixels in the 3D model. Spatial feature parameters can include the angle between the normal and the viewpoint of each pixel, the pixel's depth, etc., to determine the correspondence between pixels on the 3D model surface and their positions drawn in 2D space. This allows for stylization based on 2D dimensions, and the stylized texture features are then restored to their corresponding positions on the 3D model surface. Optionally, the spatial feature parameters of the pixels can be obtained by the renderer.
[0105] Based on the above, the method for stylizing the three-dimensional model is determined according to any of the above embodiments, based on the preset network, the stylized target image, and the spatial feature parameters of the pixels in the three-dimensional model.
[0106] The 3D model stylization method in this embodiment can automatically build a corresponding 3D model based on the 2D input images of the target to be modeled from different perspectives, and stylize the 3D model by utilizing the spatial feature parameters of pixels in the 3D model. It can model any target, meet the 3D model stylization needs of different users, and has wide applicability.
[0107] Example 4
[0108] Figure 6 This is a schematic diagram of the three-dimensional model stylization device in Embodiment 4 of this disclosure. For details not covered in this embodiment, please refer to the above embodiments.
[0109] like Figure 6 As shown, the device includes:
[0110] The acquisition module 410 is used to acquire the 3D model to be stylized and the target image for stylization.
[0111] The stylization module 420 is used to render the three-dimensional model through a preset network to obtain a two-dimensional rendered image and spatial feature parameters of pixels, and to stylize the texture features in the two-dimensional rendered image according to the spatial feature parameters and the stylization target image to obtain a stylized three-dimensional model.
[0112] The three-dimensional model stylization device in this embodiment renders the three-dimensional model into a two-dimensional image and considers the spatial relationship of each pixel. All pixels in the two-dimensional rendered image, including adjacent pixels with discontinuous texture features, can be stylized, ensuring the consistency of the three-dimensional model in spatial structure before and after stylization.
[0113] Based on the above, the stylization module 420 is specifically used for:
[0114] The field of view is determined to cover multiple perspectives of the three-dimensional model;
[0115] The three-dimensional model is rendered from the current viewpoint through the preset network to obtain the two-dimensional rendered image corresponding to the current viewpoint and the spatial feature parameters of the corresponding pixels. Based on the spatial feature parameters of the corresponding pixels of the current viewpoint and the stylized target image, the texture features in the two-dimensional rendered image corresponding to the current viewpoint are stylized.
[0116] Continue selecting the next viewpoint as the current viewpoint, and repeat the rendering and stylization operations for the current viewpoint until a complete stylized 3D model is obtained.
[0117] Based on the above, the field of view of adjacent viewpoints overlaps;
[0118] Stylize the texture features in the 2D rendered image corresponding to the current viewpoint, including:
[0119] Stylize the texture features of the parts in the 2D rendered image corresponding to the current viewpoint that did not appear in the 2D rendered image corresponding to the previous viewpoint.
[0120] Based on the above, the preset network includes a renderer and a stylization network;
[0121] The spatial feature parameters of the corresponding pixel in the current viewpoint include the angle between the normal of the corresponding pixel in the current viewpoint and the line of sight, and the depth of the corresponding pixel in the current viewpoint.
[0122] If the current viewpoint is not the first viewpoint, then the spatial feature parameters of the corresponding pixels of the current viewpoint also include the mask of the part of the two-dimensional rendered image corresponding to the current viewpoint that has been stylized in the previous viewpoint.
[0123] Based on the above, the device also includes:
[0124] The training module is used to train the preset network based on the sample model and the stylized target image until the value of the loss function of the preset network meets the requirements; wherein, the value of the loss function is determined based on the values of the following three functions:
[0125] A content loss function is used to evaluate the loss between the stylization result and the sample model;
[0126] A style loss function is used to evaluate the loss between the stylization result and the stylization target image;
[0127] A continuity loss function is used to evaluate the super-resolution test sequence VGG loss between the stylized result and the sample model whose previous viewpoint was stylized.
[0128] Based on the above, the style loss function is a weighted sum of the losses between the image features of each pixel in the stylization result and the image features of the stylized target image, wherein the weight corresponding to each pixel is negatively correlated with the angle between the normal of each pixel and the line of sight.
[0129] Based on the above, the image features of each pixel in the stylization result are determined by performing weighted convolution on the surrounding pixels of each pixel, wherein the weight of the convolution is negatively correlated with the depth difference between the pixel and the surrounding pixels.
[0130] Based on the above, the continuity loss function is the weighted sum of the VGG loss between each pixel in the stylization result and the corresponding pixel of the sample model that was stylized in the previous viewpoint, wherein the weight of each pixel is positively correlated with the angle between the normal of each pixel and the line of sight.
[0131] The above-described 3D model stylization apparatus can execute the 3D model stylization method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of the method execution.
[0132] Example 5
[0133] Figure 7 This is a schematic diagram of the three-dimensional model stylization device in Embodiment 5 of this disclosure. For details not covered in this embodiment, please refer to the above embodiments.
[0134] like Figure 7 As shown, the device includes:
[0135] Scanning module 510 is used to scan at least two two-dimensional input images, each of which includes features of the target to be modeled from a corresponding viewpoint;
[0136] The modeling module 520 is used to establish a three-dimensional model of the target to be modeled based on each of the two-dimensional input images;
[0137] The execution module 530 is used to stylize the three-dimensional model according to the preset network, the stylized target image, and the spatial feature parameters of the pixels in the three-dimensional model.
[0138] The 3D model stylization device in this embodiment uses the outline information of the first instance to guide the user to import 3D model stylization materials, thereby improving the consistency between the outlines of instances in the 3D model stylization materials and the template materials, and thus realizing the synthesis of the second instance and the background of the template material instance, improving the accuracy of 3D model stylization.
[0139] Based on the above, the method for stylizing the 3D model according to the preset network, the stylized target image, and the spatial feature parameters of the pixels in the 3D model can be determined according to the method in any of the above embodiments.
[0140] Based on the above, the structure of the execution module 530 can be found in any of the above embodiments. For example, the execution module 530 may include:
[0141] The acquisition module is used to acquire the 3D model to be stylized and the target image for stylization.
[0142] The stylization module is used to render the 3D model through a preset network to obtain a 2D rendered image and spatial feature parameters of pixels, and to stylize the texture features in the 2D rendered image according to the spatial feature parameters and the stylization target image to obtain a stylized 3D model.
[0143] The above-described 3D model stylization apparatus can execute the 3D model stylization method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of the method execution.
[0144] Example 6
[0145] Figure 8 This is a schematic diagram of the hardware structure of the electronic device in Embodiment 5 of this disclosure. Figure 8 A schematic diagram of the structure of an electronic device 600 suitable for implementing embodiments of the present disclosure is shown. The electronic device 600 in the embodiments of the present disclosure includes, but is not limited to, devices with image processing capabilities such as computers, laptops, servers, tablets, or smartphones. Figure 8 The electronic device 600 shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments disclosed herein.
[0146] like Figure 8 As shown, electronic device 600 may include one or more processing devices (e.g., central processing unit, graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. One or more processing devices 601 implement the traffic packet forwarding method provided in this disclosure. Various programs and data required for the operation of electronic device 600 are also stored in RAM 603. Processing devices 601, ROM 602, and RAM 603 are interconnected via bus 605. Input / output (I / O) interface 604 is also connected to bus 605.
[0147] Typically, the following devices can be connected to I / O interface 604: input devices 606 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 607 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 608 including, for example, magnetic tape, hard disk, etc., for storing one or more programs; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although FIG10 shows electronic device 600 with various devices, it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or included alternatively.
[0148] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, it performs the functions defined in the methods of embodiments of this disclosure.
[0149] It should be noted that the computer-readable medium described in this disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0150] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0151] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0152] The aforementioned computer-readable medium carries one or more programs. When the electronic device executes one or more of these programs, the electronic device causes the following actions: It acquires a 3D model to be stylized and a target image for stylization; it renders the 3D model using a preset network to obtain a 2D rendered image and spatial feature parameters of the pixels; and, based on the spatial feature parameters and the target image for stylization, it stylizes the texture features in the 2D rendered image to obtain a stylized 3D model. Alternatively, the electronic device causes the following actions: It scans at least two 2D input images, each of which includes features of the target to be modeled from a corresponding viewpoint; it establishes a 3D model of the target to be modeled based on each of the 2D input images; and it stylizes the 3D model based on the preset network, the target image for stylization, and the spatial feature parameters of the pixels in the 3D model.
[0153] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0154] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0155] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".
[0156] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0157] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0158] According to one or more embodiments of this disclosure, Example 1 provides a three-dimensional model stylization method, including:
[0159] Obtain the 3D model to be stylized and the target image for stylization;
[0160] The three-dimensional model is rendered using a preset network to obtain a two-dimensional rendered image and spatial feature parameters of pixels. Based on the spatial feature parameters and the stylized target image, the texture features in the two-dimensional rendered image are stylized to obtain a stylized three-dimensional model.
[0161] Example 2, following the method described in Example 1, renders the 3D model using a preset network to obtain a 2D rendered image and spatial feature parameters of pixels. Based on the spatial feature parameters and the stylized target image, the texture features in the 2D rendered image are stylized, including:
[0162] The field of view is determined to cover multiple perspectives of the three-dimensional model;
[0163] The three-dimensional model is rendered from the current viewpoint through the preset network to obtain the two-dimensional rendered image corresponding to the current viewpoint and the spatial feature parameters of the corresponding pixels. Based on the spatial feature parameters of the corresponding pixels of the current viewpoint and the stylized target image, the texture features in the two-dimensional rendered image corresponding to the current viewpoint are stylized.
[0164] Continue selecting the next viewpoint as the current viewpoint, and repeat the rendering and stylization operations for the current viewpoint until a complete stylized 3D model is obtained.
[0165] Example 3: According to the method described in Example 2, the field of view of adjacent viewpoints overlaps;
[0166] Stylize the texture features in the 2D rendered image corresponding to the current viewpoint, including:
[0167] Stylize the texture features of the parts in the 2D rendered image corresponding to the current viewpoint that did not appear in the 2D rendered image corresponding to the previous viewpoint.
[0168] Example 4: According to the method described in Example 2, the preset network includes a renderer and a stylization network;
[0169] The spatial feature parameters of the corresponding pixel in the current viewpoint include the angle between the normal of the corresponding pixel in the current viewpoint and the line of sight, and the depth of the corresponding pixel in the current viewpoint.
[0170] If the current viewpoint is not the first viewpoint, then the spatial feature parameters of the corresponding pixels of the current viewpoint also include the mask of the part of the two-dimensional rendered image corresponding to the current viewpoint that has been stylized in the previous viewpoint.
[0171] Example 5, based on the method described in Example 2, further includes:
[0172] The preset network is trained based on the sample model and the stylized target image until the value of the loss function of the preset network meets the requirements; wherein, the value of the loss function is determined based on the values of the following three functions:
[0173] A content loss function is used to evaluate the loss between the stylization result and the sample model;
[0174] A style loss function is used to evaluate the loss between the stylization result and the stylization target image;
[0175] A continuity loss function is used to evaluate the super-resolution test sequence VGG loss between the stylized result and the sample model whose previous viewpoint was stylized.
[0176] Example 6: According to the method described in Example 5, the style loss function is a weighted sum of the losses between the image features of each pixel in the stylization result and the image features of the stylized target image, wherein the weight corresponding to each pixel is negatively correlated with the angle between the normal corresponding to each pixel and the line of sight.
[0177] Example 7: According to the method described in Example 6, the image features of each pixel in the stylization result are determined by performing a weighted convolution on the surrounding pixels of each pixel, wherein the weight of the convolution is negatively correlated with the depth difference between the pixel and its surrounding pixels.
[0178] Example 8: According to the method described in Example 5, the continuity loss function is the weighted sum of the VGG loss between each pixel in the stylization result and the corresponding pixel of the sample model stylized in the previous viewpoint, wherein the weight corresponding to each pixel is positively correlated with the angle between the normal corresponding to each pixel and the line of sight.
[0179] According to one or more embodiments of this disclosure, Example 9 provides a three-dimensional model stylization method, including:
[0180] Scan at least two two-dimensional input images, each of which includes features of the target to be modeled from a corresponding viewpoint;
[0181] A three-dimensional model of the target to be modeled is established based on each of the two-dimensional input images;
[0182] The three-dimensional model is stylized based on the preset network, the stylized target image, and the spatial feature parameters of the pixels in the three-dimensional model.
[0183] According to one or more embodiments of this disclosure, Example 10 provides a three-dimensional model stylization apparatus, comprising:
[0184] The acquisition module is used to acquire the 3D model to be stylized and the target image for stylization.
[0185] The stylization module is used to render the 3D model through a preset network to obtain a 2D rendered image and spatial feature parameters of pixels, and to stylize the texture features in the 2D rendered image according to the spatial feature parameters and the stylization target image to obtain a stylized 3D model.
[0186] According to one or more embodiments of this disclosure, Example 11 provides a three-dimensional model stylization apparatus, comprising:
[0187] The scanning module is used to scan at least two two-dimensional input images, each of which includes features of the target to be modeled from a corresponding viewpoint;
[0188] The modeling module is used to establish a three-dimensional model of the target to be modeled based on each of the two-dimensional input images;
[0189] The execution module is used to stylize the 3D model according to the preset network, the stylized target image, and the spatial feature parameters of the pixels in the 3D model.
[0190] Example 12: According to the method described in Example 11, the method for stylizing the three-dimensional model based on the preset network, the stylized target image, and the spatial feature parameters of the pixels in the three-dimensional model is determined according to any one of Examples 1-8.
[0191] According to one or more embodiments of this disclosure, Example 13 provides an electronic device comprising:
[0192] One or more processors;
[0193] Storage device for storing one or more programs;
[0194] When the one or more programs are executed by the one or more processors, the one or more processors implement the three-dimensional model stylization method as described in any of Examples 1-10.
[0195] According to one or more embodiments of this disclosure, Example 14 provides a method for stylizing a three-dimensional model as described in any of Examples 1-10 when the program is executed by a processor.
[0196] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0197] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0198] Although this topic has been described using language specific to structural features and / or methodological logic, it should be understood that the topic defined in the accompanying example book is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the example book.
Claims
1. A method for stylizing a 3D model, comprising: Obtain the 3D model to be stylized and the target image for stylization; The field of view is determined to cover multiple perspectives of the three-dimensional model; The three-dimensional model is rendered from the current viewpoint through a preset network to obtain the two-dimensional rendered image corresponding to the current viewpoint and the spatial feature parameters of the corresponding pixels. Based on the spatial feature parameters of the corresponding pixels of the current viewpoint and the stylized target image, the texture features in the two-dimensional rendered image corresponding to the current viewpoint are stylized. Continue selecting the next viewpoint as the current viewpoint, and repeat the rendering and stylization operations for the current viewpoint until a complete stylized 3D model is obtained; The preset network includes a renderer and a stylization network; The spatial feature parameters of the corresponding pixel in the current viewpoint include the angle between the normal of the corresponding pixel in the current viewpoint and the line of sight, and the depth of the corresponding pixel in the current viewpoint. If the current viewpoint is not the first viewpoint, then the spatial feature parameters of the corresponding pixels of the current viewpoint also include the mask of the part of the two-dimensional rendered image corresponding to the current viewpoint that has been stylized in the previous viewpoint.
2. The method of claim 1, wherein, The field of view of adjacent viewpoints overlaps; Stylize the texture features in the 2D rendered image corresponding to the current viewpoint, including: Stylize the texture features of the parts in the 2D rendered image corresponding to the current viewpoint that did not appear in the 2D rendered image corresponding to the previous viewpoint.
3. The method of claim 1, wherein, Also includes: The preset network is trained based on the sample model and the stylized target image until the value of the loss function of the preset network meets the requirements; wherein, the value of the loss function is determined based on the values of the following three functions: A content loss function is used to evaluate the loss between the stylization result and the sample model; A style loss function is used to evaluate the loss between the stylization result and the stylization target image; A continuity loss function is used to evaluate the super-resolution test sequence VGG loss between the stylized result and the sample model whose previous viewpoint was stylized.
4. The method of claim 3, wherein, The style loss function is a weighted sum of the losses between the image features of each pixel in the stylization result and the image features of the stylized target image, wherein the weight of each pixel is negatively correlated with the angle between the normal of each pixel and the line of sight.
5. The method of claim 4, wherein, The image features of each pixel in the stylization result are determined by performing weighted convolution on the surrounding pixels of each pixel, wherein the weight of the convolution is negatively correlated with the depth difference between the pixel and the surrounding pixels.
6. The method of claim 3, wherein, The continuity loss function is the weighted sum of the VGG loss between each pixel in the stylization result and the corresponding pixel in the sample model of the previous viewpoint that was stylized. The weight of each pixel is positively correlated with the angle between the normal of each pixel and the line of sight.
7. A method of stylizing a three-dimensional model, the method comprising: include: Scan at least two two-dimensional input images, each of which includes features of the target to be modeled from a corresponding viewpoint; A three-dimensional model of the target to be modeled is established based on each of the two-dimensional input images; Based on a preset network, a stylized target image, and spatial feature parameters of pixels in the 3D model, the 3D model is stylized according to any one of claims 1-6.
8. A three-dimensional model stylization apparatus, characterized by: include: The acquisition module is used to acquire the 3D model to be stylized and the target image for stylization. The stylization module is used to render the three-dimensional model through a preset network to obtain a two-dimensional rendered image and spatial feature parameters of pixels, and to stylize the texture features in the two-dimensional rendered image according to the spatial feature parameters and the stylization target image to obtain a stylized three-dimensional model. The stylization module is specifically used for: The field of view is determined to cover multiple perspectives of the three-dimensional model; The three-dimensional model is rendered from the current viewpoint through a preset network to obtain the two-dimensional rendered image corresponding to the current viewpoint and the spatial feature parameters of the corresponding pixels. Based on the spatial feature parameters of the corresponding pixels of the current viewpoint and the stylized target image, the texture features in the two-dimensional rendered image corresponding to the current viewpoint are stylized. Continue selecting the next viewpoint as the current viewpoint, and repeat the rendering and stylization operations for the current viewpoint until a complete stylized 3D model is obtained; The preset network includes a renderer and a stylization network; The spatial feature parameters of the corresponding pixel in the current viewpoint include the angle between the normal of the corresponding pixel in the current viewpoint and the line of sight, and the depth of the corresponding pixel in the current viewpoint. If the current viewpoint is not the first viewpoint, then the spatial feature parameters of the corresponding pixels of the current viewpoint also include the mask of the part of the two-dimensional rendered image corresponding to the current viewpoint that has been stylized in the previous viewpoint.
9. A three-dimensional model stylization apparatus, comprising: include: The scanning module is used to scan at least two two-dimensional input images, each of which includes features of the target to be modeled from a corresponding viewpoint; The modeling module is used to establish a three-dimensional model of the target to be modeled based on each of the two-dimensional input images; An execution module is used to stylize the three-dimensional model according to any one of claims 1-6, based on a preset network, a stylized target image, and spatial feature parameters of pixels in the three-dimensional model.
10. An electronic device, comprising: include: One or more processors; Storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the three-dimensional model stylization method as described in any one of claims 1-7.
11. A computer readable storage medium having stored thereon a computer program, characterized in that When the program is executed by the processor, it implements the three-dimensional model stylization method as described in any one of claims 1-7.
Citation Information
Patent Citations
Image Style Transfer for Three-Dimensional Models
CN110084874A