An Optimization Method for Example-based 3D Texture Procedural Generation Model

The integration of program rule-based and example-based texture generation with spatial tags optimizes the texture creation process, enhancing efficiency and flexibility for artists by automating texture generation and application across scenes.

CN114708379BActive Publication Date: 2025-07-15ZHEJIANG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210176364.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-25
Publication Date
2025-07-15
Estimated Expiration
2042-02-25

AI Technical Summary

Technical Problem

The existing three-dimensional texture generation method based on examples is highly automated but difficult to edit. The writing of the method based on program rules is complex and time-consuming. Users need to adjust the node parameters a lot, making it difficult to efficiently generate textures that meet expectations.

Method used

Combining node graphs and microrenderable technology, the texture generation model is optimized through style loss functions, spatial labels are introduced to reduce the interference of the rendering process, features are extracted using convolutional neural networks, and parameters are adjusted using gradient optimization algorithms.

Benefits of technology

It improves the efficiency and flexibility of texture generation, assists art workers in quickly creating three-dimensional textures of specific styles, reduces manpower investment, and improves production efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114708379B_ABST
    Figure CN114708379B_ABST
Patent Text Reader

Abstract

The present invention discloses an optimization method for an example-based three-dimensional texture procedural generation model. Based on the procedural texture generation technology based on program rules, this method automatically optimizes the parameters of the generation model according to example images. The present invention uses differentiable rendering technology to convert the generated texture into an image representing the appearance of an object, and optimizes the parameters of the generation model with the style loss between this image and the example image. Moreover, spatial tags are added to the loss function to reduce the interference of the ambiguity brought by the rendering process on the optimization effect. Compared with the existing procedural texture generation technology based on program rules, this method has a higher degree of automation; while compared with the existing example-based generation technology, this method allows users to edit the generation model, making it more flexible. In addition, this method is the first parameter optimization method for the procedural generation model of three-dimensional textures.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer graphics, and particularly to an optimization method for an example-based three-dimensional texture process generation model. Background Art

[0002] In computer graphics, texture is a common method for representing high-frequency detail information in space. By wrapping the texture on the surface of an object, attributes such as the undulation height, material color, and material roughness at each point on its surface can be described; by mapping a three-dimensional texture to a specific area in three-dimensional space, specific attributes of this area can be described, such as the light scattering rate at each position in the cloud or the albedo at each position in a translucent material.

[0003] How to efficiently obtain texture data has always been a much-concerned issue. Manual creation has high labor and time costs, and capturing real-world texture data is also easily limited in terms of scale and efficiency. Therefore, when an application requires a large number of textures with certain statistical characteristics, the process generation method is often a suitable choice. Such textures are widely present in life, including both textures representing colors and textures representing material attributes such as normal vectors, undulation heights, and medium densities, such as the texture color of the trunk of a tree and the normal details of the surface of a brick wall. Texture process generation methods can be roughly divided into two categories:

[0004] I. Example-based process generation. The so-called example-based means that the user provides a two-dimensional / three-dimensional texture as an example, and the generation algorithm batch-produces textures with a similar visual style to the example image. Such methods can be further divided into stitching algorithms based on local similarity, algorithms based on texture optimization, and end-to-end methods based on neural networks.

[0005] II. Program rule-based process generation, that is, a generation model for generating texture data is designed manually, which contains complete program code. Due to the high learning cost of writing program code, visual programming tools based on node graphs are very popular among designers. In these tools, designers can construct programs by combining nodes with specific functions.

[0006] These two types of methods each have their own advantages and disadvantages. Example-based methods are highly automated. Usually, users only need to provide an example image to obtain a large number of target textures similar to it. However, the intermediate processes of these methods are almost like black boxes and cannot be edited. In production practice, users often need to observe the results frequently and adjust the generation process, and this requirement limits the application scope of example-based methods. In contrast, methods based on procedural rules are easy to understand. Users can edit the generation rules and modify certain features of the target texture. The disadvantage of this type of method is that it is not easy to write the generation rules. Even for experienced artists, it often takes a lot of time to design the node graph and adjust the node parameters to generate textures with specific visual features.

[0007] Writing a texture generation program requires a certain programming foundation, and it is also difficult to observe the intermediate results during the writing process, which is very inconvenient for artists and art workers. To lower the threshold of writing procedural rules and improve the writing efficiency, a visual programming method based on node graphs is introduced into this field. Currently, most commercial software for procedural generation uses this programming mode, such as Houdini, Substance Painter, etc. To make the generation rules written in the node graph produce textures that meet expectations, experienced artists / art workers are required to design the structure of the node graph. Even with the node graph itself, different node parameters will produce completely different target textures. Therefore, even though there are hundreds or thousands of node graphs accumulated in commercial software that can generate various textures, users still need to spend a lot of time adjusting the node parameters, which goes against the original intention of the procedural generation method of "efficiently generating textures". Summary of the Invention

[0008] The purpose of the present invention is to provide an optimization method for an example-based three-dimensional texture procedural generation model in view of the deficiencies of the prior art.

[0009] The purpose of the present invention is achieved through the following technical solutions: An optimization method for an example-based three-dimensional texture procedural generation model, comprising the following steps:

[0010] Step 1: The user inputs a texture procedural generation model based on procedural rules in the form of a node graph, combines the three-dimensional target texture output by this model with the pre-prepared scene data, and obtains an image representing the appearance of the object through differentiable rendering.

[0011] Step 2: Input the appearance image obtained in Step 1 and the example image provided by the user into the style loss function, and numerically optimize the parameters in the generation model based on the value of this loss function.

[0012] Step 3: Add spatial labels to the style loss function in Step 2 to reduce the interference of the ambiguity brought by the rendering process on the optimization effect.

[0013] Furthermore, Step 1 is implemented through the following sub-steps:

[0014] (1.1) Use a node graph to describe the process generation model of 3D textures, and use code generation technology to efficiently evaluate and differentiate the 3D texture node graph.

[0015] (1.2) Render the generated texture into an image by means of a differentiable rendering technology applicable to 3D texture data with high-order scattering and high resolution.

[0016] (1.3) Introduce a path caching technology to improve the problem that the traditional differentiable rendering technology requires a large amount of storage space when the number of scattering times is high. By traversing the same ray path multiple times, the required storage space is significantly reduced.

[0017] Furthermore, Step 2 is implemented through the following sub-steps:

[0018] (2.1) Input the appearance image and the example image into a convolutional neural network for computer vision tasks respectively, and extract the key feature layers therein.

[0019] (2.2) For the two feature maps corresponding to the appearance image and the example image in a certain feature layer respectively, erase the spatial position information of all their feature pixels to form two lists of feature samples containing M N-dimensional vectors, where M is the number of feature image pixels and N is the number of feature pixel channels.

[0020] (2.3) Randomly select a certain number of direction vectors in the N-dimensional space. For each direction vector, calculate the projection length of each feature sample in the two lists in (2.2) on this direction vector to obtain two projection vectors with a length of M.

[0021] (2.4) Sort the elements of the two projection vectors separately from smallest to largest.

[0022] (2.5) Calculate the Euclidean distance between the two projection vectors, divide it by the total number of direction vectors, and accumulate it to the final loss function value.

[0023] Furthermore, the specific problems and solutions in Step 3 are:

[0024] During rendering, factors such as the shape of objects in the scene and the ambient light source can cause color differences in different regions of the objects in the appearance image. This part of the difference is inherent in the scene. If the evaluation function attributes this part of the difference to the target texture and then to the influence of the node graph parameters, it will interfere with the optimization result. To solve this problem, the present invention introduces the concept of spatial tags. A spatial tag is an image with the same resolution as the appearance image and the example image, used to describe the inherent differences between different regions on the appearance image. The values of spatial tags are flexible and can be any quantity that can reflect the inherent differences. For example, using the world space normal as the value of the spatial tag is equivalent to regarding the normal difference as an inherent difference. If there are obvious differences in the normals of region A and region B on the image, then the styles between A and B should not be compared. Specifically, the detailed steps for generating spatial tags are as follows:

[0025] (3.1) Replace the target texture in the scene data with a constant texture, use differentiable rendering to convert the scene into an appearance image, and calculate the sum of squared differences between the appearance image and the example image.

[0026] (3.2) Perform gradient descent optimization on the texture based on the gradient of the texture value with respect to the difference calculated in (3.1), and use the appearance image corresponding to the constant texture optimized to convergence as the spatial tag.

[0027] (3.3) For all feature maps obtained in (2.1) with the same resolution as the spatial tag, append the spatial tag as an additional feature channel after the original feature channels.

[0028] (3.4) For each randomly selected direction vector (v1, v2, …, v N ), replace it with D direction vectors {(v1, …, v N , 0, …, a m , …, 0)|m = 1, …, D}, where D is the number of channels of the spatial tag generated in (3.2).

[0029] The beneficial effects of the present invention are as follows: The present invention describes the texture generation model with program rules and then optimizes the model parameters with examples, combining the advantages of these two types of technologies. It saves manpower compared to the former and is more flexible in editing compared to the latter. The present invention can assist art workers in quickly creating three-dimensional textures with specific styles, improving the production efficiency of texture data, and transferring the texture of a specific object on an image to other rendered objects. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 is a flowchart of an optimization method for an example-based three-dimensional texture process generation model;

[0031] Figure 2It is a diagram of the core system components required to implement this method. Detailed implementation manners

[0032] The present invention will be described in detail below with reference to the accompanying drawings.

[0033] As Figure 1 and Figure 2 shown, the method for optimizing an example-based 3D texture process generation model according to the present invention includes the following steps:

[0034] The object of the present invention is achieved by the following technical solutions: A method for optimizing an example-based 3D texture process generation model, including the following steps:

[0035] Step 1: The user inputs a texture process generation model based on program rules in the form of a node graph, combines the 3D target texture output by this model with the pre-prepared scene data, and obtains an image representing the appearance of the object through differentiable rendering.

[0036] This step is the basic step for generating the target texture of the present invention and is divided into the following sub-steps.

[0037] 1) Parse the node graph structure input by the user to generate a function for evaluating the node graph and a function for taking the derivative. The algorithm for generating the evaluation function is as follows:

[0038] 1 symbol generate_eval(ni){

[0039] 2 V = empty list

[0040] 3 For each input node nj of node ni {

[0041] 4 v = generate_eval(nj)

[0042] 5 V = append(V, v)

[0043] 6}

[0044] 7 r = create a new variable

[0045] 8 Generate code for calculating r from V according to the type of ni

[0046] 9 return r

[0047] 10}

[0048] The algorithm for generating the derivative function is as follows:

[0049] 1 symbol generate_grad(ni, dLdni){

[0050] 2 Generate the code to calculate the partial derivative of \(n_i\) with respect to the parameter \(\theta\), and store the result in the variable \(c\).

[0051] 3 Generate the code: \(dLd\theta=dLd\theta + dLd n_i*c\).

[0052] 4 For each input node \(n_j\) of the node \(n_i\) {

[0053] 5 Generate the code to calculate the partial derivative of \(n_j\) with respect to the parameter \(\theta\), and store the result in the variable \(d\).

[0054] 6 Generate the code: \(d = d * dLd n_i\).

[0055] 7 generate_grad(\(n_j\), \(d\))

[0056] 8}

[0057] 9}

[0058] 2) Use differentiable rendering to render the target texture obtained from the evaluation node graph into an appearance image, and at the same time calculate the derivative of the appearance image with respect to the target texture. Among them, the forward rendering process can adopt the relatively mature path tracing algorithm. When calculating the derivative in reverse, for each ray path starting from the camera, it is necessary to traverse the path once first, record the coefficients at each vertex, and after the traversal is completed, use these coefficients to calculate the radiance estimate from the next vertex at each vertex; then traverse the path again, substitute the previously calculated radiance estimate, and accumulate to obtain the derivative of the appearance image with respect to the target texture.

[0059] Step 2: Input the appearance image obtained in Step 1 and the example image provided by the user into the style loss function together, and numerically optimize the parameters in the generation model based on the value of this loss function.

[0060] This step is the core of the present invention and is divided into the following sub-steps.

[0061] 1) Input the appearance image and the example image into the convolutional neural network for computer vision tasks respectively, and extract the key feature layers therein. Here, either the VGG-16 or VGG-19 network can be used for the neural network.

[0062] 2) For the two feature maps corresponding to the appearance image and the example image in a certain feature layer respectively, erase the spatial position information of all their feature pixels to form two feature sample lists containing \(M\) \(N\)-dimensional vectors, where \(M\) is the number of feature image pixels and \(N\) is the number of feature pixel channels.

[0063] 3) Randomly select a certain number of direction vectors in the N-dimensional space. For each direction vector, calculate the projection lengths of each feature sample in the two lists on this direction vector to obtain two projection vectors of length M.

[0064] 4) For each direction vector, sort the elements of the two corresponding projection vectors from smallest to largest.

[0065] 5) For each direction vector, calculate the Euclidean distance between the two corresponding projection vectors, divide it by the total number of direction vectors, and accumulate it to the final loss function value.

[0066] After obtaining the loss function value, use the gradient-guided Markov chain Monte Carlo algorithm, i.e., the Metropolis-Adjusted Langevin algorithm, to sample the node values.

[0067] Step 3: Add spatial labels to the style loss function in Step 2 to reduce the interference of the ambiguity brought by the rendering process on the optimization effect.

[0068] This step plays an important role in the effect of the present invention in complex scenarios and is divided into the following sub-steps.

[0069] 1) Replace the target texture in the scene data with a constant texture, use differentiable rendering to convert the scene into an appearance image, and calculate the sum of squared differences between the appearance image and the example image.

[0070] 2) Based on the gradient of the texture value with respect to the difference calculated in (1), perform gradient descent optimization on the texture, and use the appearance image corresponding to the constant texture optimized to convergence as the spatial label.

[0071] 3) For all feature maps with the same resolution as the spatial label obtained in the previous step, append the spatial label as an additional feature channel after the original feature channels.

[0072] 4) For each direction vector (v1, v2, …, v N ), randomly selected when calculating the loss function, replace it with D direction vectors {(v1, …, v N , 0, …, a m , …, 0)|m = 1, …, D}, where D is the number of channels of the spatial label generated in (3.2).

Claims

1. An optimization method for an example-based 3D texture procedural generation model, characterized in that It includes the following steps: Step 1: Input a texture process generation model based on program rules in the form of a node graph, combine the three-dimensional target texture output by this model with the pre-prepared scene data, and obtain an image representing the object appearance through differentiable rendering; Step 2: Input the appearance image obtained in Step 1 and the example image into a style loss function together, and numerically optimize the parameters in the generation model based on the value of this loss function; The implementation of Step 2 is achieved through the following sub-steps: (2.1) Input the appearance image and the example image into a convolutional neural network for computer vision tasks respectively, and extract the key feature layers therein; (2.2) For the two feature maps corresponding to the appearance image and the example image in a certain feature layer respectively, erase the spatial position information of all their feature pixels, forming two feature sample lists containing M N-dimensional vectors, where M is the number of feature image pixels and N is the number of feature pixel channels; (2.3) Randomly select a certain number of direction vectors in the N-dimensional space; for each direction vector, calculate the projection length of each feature sample in the two lists in (2.2) on this direction vector, obtaining two projection vectors with a length of M; (2.4) Sort the elements of the two projection vectors separately from small to large; (2.5) Calculate the Euclidean distance between the two projection vectors, divide it by the total number of direction vectors, and accumulate it to the final loss function value; Step 3: Add a spatial label to the style loss function in Step 2 to reduce the interference of the ambiguity brought by the rendering process on the optimization effect; The implementation of Step 3 is achieved through the following sub-steps: (3.1) Replace the target texture in the scene data with a constant texture, use differentiable rendering to convert the scene into an appearance image, and calculate the sum of squares difference between the appearance image and the example image; (3.2) Perform gradient descent optimization on the texture based on the gradient of the texture value with respect to the difference calculated in (3.1), and use the appearance image corresponding to the constant texture optimized to convergence as the spatial label; (3.3) For all feature maps obtained in (2.1) with the same resolution as the spatial label, append the spatial label as an additional feature channel after the original feature channel; (3.4) For each direction vector (v1, v2, …, v N ), randomly selected in (2.3), replace it with N+D direction vectors {(v1, …, v N , 0, …, α m , …, 0)| m = 1, …, D}, where D is the number of channels of the spatial labels generated in (3.2).

2. The method for optimizing a three-dimensional texture process generation model based on examples according to claim 1, wherein The specific content of Step 1 is as follows: (1.1) Use a node graph to describe the process generation model of three-dimensional texture, and use code generation technology to evaluate and differentiate the three-dimensional texture node graph; (1.2) Use the Monte Carlo method to trace the light transmission path in the scene, so as to render the generated texture into an image; for each light path in the scene, continuously randomly sample the scattering direction on the unit sphere, starting from the camera, and continuously use ray tracing to find the next scattering point based on the current scattering point position. During this process, accumulate the contribution of the light source at each scattering point to the radiance carried by this transmission path; (1.3) Use the reverse derivative method to calculate the derivative of the image rendered in (1.2) with respect to the texture; For each light path starting from the camera, first traverse this path once, record the coefficients at each vertex, and after the traversal, use these coefficients to calculate the radiance estimation amount from the next vertex at each vertex; Then traverse this path again, substitute the previously calculated radiance into the chain rule of differentiation, and accumulate it to obtain the derivative of the rendered image with respect to the texture.

Citation Information

Patent Citations

  • Generating procedural materials from digital images

    US20210343051A1