Differential rendering and deep-learning based generation of downsampled textured maps

By applying geometric domain filters and warping UV mappings, the method addresses the distortion issues in existing texture map downsampling techniques, achieving high-fidelity, low-resolution texture maps that reduce computational requirements.

WO2025106646A1PCT designated stage expired Publication Date: 2025-05-22DOLBY LABORATORIES LICENSING CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/055882
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-11
Filing Date
2024-11-14
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

Existing methods for downsampling 3D texture maps independently treat texture map downsampling and 3D mesh decimation, failing to account for the geometry of the texture map, leading to distortions when mapped to 3D models.

Method used

The proposed solution involves applying filters in the geometric domain that are compliant with the original surface topology of the texture map and its UV mapping, using deep learning activations and graph filters to distribute filtering weights based on geometric and textural importance, while also warping UV mapping to reduce distortions.

Benefits of technology

This approach effectively reduces the resolution of texture maps while retaining the post-rendering appearance of the original high-resolution version, improving visual fidelity and reducing computational resources required for rendering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024055882_22052025_PF_FP_ABST
    Figure US2024055882_22052025_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods for generating downsampled textured maps. One example method includes training a neural network by providing a three-dimensional model including an original texture map to the neural network, downsampling the original texture map, and iteratively differentiably rendering the original and downsampled texture maps and using differences in renderings of the original and downsampled texture maps as feedback for training the neural network. After reaching a training completion condition, the trained neural network provides a downsampled texture map that has a lower resolution than the texture map.
Need to check novelty before this filing date? Find Prior Art

Description

DIFFERENTIAL RENDERING AND DEEP-LEARNING BASED GENERATION OF DOWNSAMPLED TEXTURED MAPS1. Cross-Reference to Related Applications

[0001] This application claims the benefit of priority from US Provisional Application Ser. No. 63 / 600,166, filed on 17 November 2023, and European Patent Application No. EP 24151481.9, filed on 11 January 2024, each of which is incorporated by reference herein in its entirety.2. Field of the Disclosure

[0002] This application relates generally to systems and methods of generating downsampled 3D textures for computer, mobile, and virtual reality graphics applications.3. Background

[0003] Three-dimensional (3D) mesh models are defined by meshes and texture maps. The meshes are a collection of faces that are constituted by vertices and edges. Geometrical faces are commonly represented by polygons that, when combined, form surfaces of the 3D model. Texture maps are a two-dimensional (2D) representation of the model surface properties, such as diffuse texture maps, normal maps, and the like. The texture maps are projected onto the model surfaces by means of UV-mapping that sets correspondence between the mesh vertices coordinates and the coordinates in the texture maps.

[0004] The 3D model may be, for example, individual objects, environmental surfaces, and the like. To represent the color of the 3D model, the 3D model is flattened onto a 2D image such that each triangle in the 3D model is mapped to a polygon in the 2D image. In this manner, the pixel values of the triangles in the 2D image represent the texture of the associated face on the 3D model.

[0005] Mesh models are more compact and can be efficiently rendered when compared to other representation methods, such as point clouds, voxel grids, and volumetric fields. However, to have high-fidelity models, textured meshes have large polygon counts (e.g., a large number of triangles) and can still consume excessive bandwidth, storage, and rendering resources.BRIEF SUMMARY OF THE DISCLOSURE

[0006] Highly-detailed meshes that include a large number of polygons and a high degree of color information often require a large amount of computational resources to render. However, the needed computational resources are not always available. For example, wireless virtual reality (VR) headsets include a local graphical processing unit (GPU) that may not be capable of rendering such high-resolution textured meshes.

[0007] In some instances, the 2D textured maps are downsampled to reduce the required computational power for rendering the associated 3D models. Known downsampling methods, such as bilinear and bicubic downsampling of texture maps, treat down-sampling of texture maps and decimation of 3D meshes independently of each other. Additionally, known downsampling methods do not account for the geometry of the texture map while downsizing, maintain the same mapping between triangles and texture coordinates while downsampling, and use the same downsizing kernel for each texture map in a scene. For example, in statistical image downsampling methods such as bicubic and lanczos, neighboring pixels in each direction of a given pixel are considered to be equidistant and an equal weight is assigned to each neighboring pixel. However, unlike regular 2D images, texture maps sit on irregular non-planar surfaces and are warped and distorted versions of the original layout. Hence, the assumption of equidistant neighbors causes distortions in the downsampled texture when mapped to a 3D model.

[0008] Embodiments described herein reduce the resolution of texture maps while more closely retaining the post-rendering appearance compared to the original high-resolution version of the texture map. Particularly, filters are applied in the geometric domain that are compliant with the original surface topology of the texture map and its UV mapping. The filtering includes sampling intermediate deep learning activations at vertex locations and processing the vertex locations using graph filters or geometric filters, ensuring the filtering weights are distributed according to geometric and textural importance. Additionally, warping of the UV mapping is performed to reduce distortions when textures are shaded onto the 3D model.

[0009] In one exemplary aspect of the present disclosure, there is provided a method for generating downsampled texture maps. The method includes training a neural network by (i) providing, to the neural network, a three-dimensional (3D) model including an original texture map that covers a 3D mesh, (ii) rendering the original texture map on the 3D mesh and rendering a downsampled texture map on the 3D mesh or on a downsampled version of the 3D mesh based on current parameters of the neural network; (iii) generating loss function feedback by analyzingdifferences between the rendered downsampled texture map and the rendered original texture map,(iv) using the loss function feedback to alter the parameters of the neural network, and (v) iterating (ii) through (iv) until a completion condition is satisfied; and generating, with the trained neural network, at least one final downsampled texture map, wherein the final downsampled texture map has a lower resolution than the original texture map.

[0010] In another exemplary aspect of the present disclosure, there is provided a system for generating downsampled textures. The system includes an electronic processor connected to a memory. The memory stores a neural network. The electronic processor is configured to train the neural network by (i) providing, to the neural network, a three-dimensional (3D) model including an original texture map that covers a 3D mesh, (ii) rendering the original texture map on the 3D mesh and rendering a downsampled texture map on the 3D mesh or on a downsampled version of the 3D mesh based on current parameters of the neural network, (iii) generating loss function feedback by analyzing differences between the rendered downsampled texture map and the rendered original texture map, (iv) using the loss function feedback to alter the parameters of the neural network, and(v) iterating (ii) through (iv) until a completion condition is satisfied; and generate, with the trained neural network, at least one final downsampled texture map, wherein the final downsampled texture map has a lower resolution than the original texture map.

[0011] In this manner, various aspects of the present disclosure provide for the display of images having a high dynamic range and high resolution, and effect improvements in at least the technical fields of graphics engines, virtual reality, signal processing, and the like.DESCRIPTION OF THE DRAWINGS

[0012] These and other more detailed and specific features of various embodiments are more fully disclosed in the following description, reference being had to the accompanying drawings, in which:

[0013] FIG. 1 illustrates an example process for a video delivery pipeline.

[0014] FIG. 2 illustrates a block diagram of an example texture map generation operation.

[0015] FIG. 3 illustrates a block diagram of an example differentiable rendering operation.

[0016] FIG. 4 illustrates a 3D model differentiably rendered from a set of random camera angles.

[0017] FIG. 5 illustrates a block diagram of another example texture map generation operation.

[0018] FIG. 6 illustrates an example of a plurality of features detected by a neural network.

[0019] FIG. 7 illustrates an example output of a sample vertex-wise features module included in the texture map generation operation of FIG. 5.

[0020] FIG. 8 illustrates an example output of a graph convolutions module included in the texture map generation operation of FIG. 5.

[0021] FIG. 9 A illustrates an example original texture map.

[0022] FIG. 9B illustrates an example downsampled texture map.

[0023] FIGS. 10A-10B illustrate an area of interest of the original texture map of FIG. 9A compared with the same area of interest of the downsampled texture map of FIG. 9B.

[0024] FIGS. 11A-1 ID are graphs illustrating quality and performance gains provided by the texture map generation operation of FIG. 5.

[0025] FIG. 12 illustrates visual improvements provided by the texture map generation operation of FIG. 5.DETAILED DESCRIPTION

[0026] This disclosure and aspects thereof can be embodied in various forms, including hardware, devices or circuits controlled by computer-implemented methods, computer program products, computer systems and networks, user interfaces, and application programming interfaces; as well as hardware-implemented methods, signal processing circuits, memory arrays, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), and the like. The foregoing is intended solely to give a general idea of various aspects of the present disclosure, and does not limit the scope of the disclosure in any way.

[0027] In the following description, numerous details are set forth, such as device configurations, timings, operations, and the like, in order to provide an understanding of one or more aspects of the present disclosure. It will be readily apparent to one skilled in the art that these specific details are merely exemplary and not intended to limit the scope of this application.

[0028] Moreover, while the present disclosure focuses mainly on examples in which the various circuits are used in digital projection systems, it will be understood that these are merely examples. It will further be understood that the disclosed systems and methods can be used in any device in which there is a need to project light; for example, cinema, consumer, and other commercial projection systems, heads-up displays, virtual reality displays, and the like. Disclosed systems and methods may be implemented in additional display devices, such as with an OLED display, an LCD display, a quantum dot display, or the like.

[0029] FIG. 1 depicts an example process of a content delivery pipeline 100, showing various stages from content capture and / or creation to content display according to an embodiment. A sequence of frames 102 may be captured or generated using a content generation block 105. The frames 102 may be digitally captured (e.g., by a digital camera) or generated by a computer (e.g., using computer animation) to provide data 107. For example, a graphics engine may be used to generate 3D models that are provided as data 107. The data 107 may include video, images, or 3D environments that are provided via a display 140.

[0030] In a production phase 110, the data 107 may be edited to provide a production stream 112. The data of the production stream 112 may then be provided to a processor (or one or more processors, such as a central processing unit, CPU) at a post-production block 115 for postproduction editing. The post-production editing of the block 115 may include, e.g., adjusting or modifying colors or brightness in particular areas of an image to enhance the image quality or achieve a particular appearance for the image in accordance with the video creator’s creative intent. This part of post-production editing is sometimes referred to as “color timing” or “color grading.” Other editing (e.g., scene selection and sequencing, image cropping, addition of computer-generated visual special effects, etc.) may be performed at the block 115 to yield a “final” version 117 of the production for distribution. For example, rendering techniques described herein, such as the generation of texture maps, may be performed during post-production editing 115. During the postproduction editing 115, the content may be viewed on a reference display 125. In other instances, the rendering and map generation techniques described herein may be performed at contentgeneration block 105, production phase 110, or another suitable processing step within the content delivery pipeline 100.

[0031] Following the post-production 115, data of the final version 117 may be delivered (as a signal) to a coding block 120 for being delivered downstream to decoding and playback devices, such as VR headsets, near-eye displays, and the like. In some embodiments, the coding block 120 may include audio and video encoders, such as those defined by the ATSC, DVB, DVD, Blu-Ray, and other delivery formats, to generate a coded bitstream 122. In a receiver, the coded bitstream 122 is decoded by a decoding unit 130 to generate a corresponding decoded signal 132 representing a copy or a close approximation of the final version 117. The receiver may be attached to a target display 140 that may have somewhat or completely different characteristics than the reference display 125. In such cases, a display management (DM) block 135 may be used to map the decoded signal 132 to the characteristics of the target display 140 by generating a display-mapped signal 137. Depending on the embodiment, the decoding unit 130 and display management block 135 may include individual processors or may be based on a single integrated processing unit. The coded bitstream 122 may include metadata to assist in the reconstruction of the final version 117.

[0032] In most interactive 3D graphics applications such as video games, virtual reality applications, and augmented reality applications, a majority of graphics assets are depicted using textured meshes. These meshes may be designed by graphics artists who build complex objects by blending primitive 3D shapes or sculpting complex geometric surfaces. There has also been an increase in 3D content captured from real world scenes in recent years. While some methods exist that may directly generate meshes from input views, most of these methods produce 3D reconstructions in the format of point clouds, voxel grids, or volumetric fields which are converted to texture maps.

[0033] A UV parametrization process may be performed to generate texture maps from textured 3D mesh surfaces. The process of UV parametrization involves unwrapping the 3D model into the two-dimensional UV space and generating the mapping that defines how each point on the 3D model corresponds to a point on the UV plane. To reduce distortions due to the flattening from 3D to 2D, the texture map may be divided into multiple patches, with the flattening and UV parametrization performed independently on each patch. However, due to noisy 3D reconstruction of geometry, the division of mesh into patches and their flattening often suffer from unnatural texture scams, inefficient utilization of the space on the UV plane and sub-optimal calculation ofUV mapping. This generation of texture maps of reconstructed 3D scenes may incur a cost in terms of wasted memory usage while rendering and reduced quality of rendered images. These problems can be exacerbated when the resolution of the texture map is reduced.

[0034] Additionally, deep learning based techniques (for example, neural networks) have been developed for analyzing textured meshes, including 3D object recognition, detection, and segmentation. Specific pooling operations have also been devised for textured meshes which allow them to retain triangular face structure and face orientation. While graph convolutions are suitable for acting on the geometry and connectivity of the mesh and regular convolutions are capable of acting upon 2D texture maps, known techniques do not have a framework for processing them jointly.

[0035] Embodiments described herein utilize a neural network to perform downsampling (e.g., downsizing) of received texture maps. For example, embodiments described herein receive a textured 3D mesh as an input and derive a texture downsampling function and the corresponding downsampled textures as an output. Embodiments described herein utilize both the geometry of the 3D model and the layout and the UV parametrization to generate a downsampled texture map that improves the visual fidelity of the rendered view.

[0036] FIG. 2 illustrates an operation 200 of generating texture maps for 3D models. An original texture map 202 is received by an electronic processor of the content delivery pipeline 100 (for example, an electronic processor of the post-production block 115). A downsampling neural network 204 receives the original texture map 202 and outputs a plurality of down-sized features 206, an example of which is shown in FIG. 6. In some instances, the downsampling neural network 204 also receives pointwise normal vectors projecting from the vertices of polynomials within the original texture map 202. The plurality of down- sized features 206 are reconstructed by a reconstruction operation 208 to generate a downsized texture map 210.

[0037] The downsampling neural network 204 is a neural network (e.g., a convolutional neural network [CNN]) that identifies important features within the received texture map 202. For example, given an input mesh M (T, v,f, u) with a texture size of H x W, the downsampling neural network 204 projects the 3D geometric properties of the mesh, specifically the vertex positions and normal, to a 2D map that aligns with the original texture map 202. The downsampling neural network 204 may utilize the UV parametrization and barycentric coordinates to smoothly interpolatethe geometric features into a 2D feature map (e.g., the down-sized features 206). As the same UV parametrization is used for projecting textures onto the surface of the mesh, the texture map aligns with the geometric feature maps. The conversion of the geometric information to the 2D domain allows it to be filtered through the use of regular convolutional layers. In some instances, the feature maps containing information of 3D properties of the mesh are concatenated with the original texture map 202 before being input to the downsampling neural network 204 (for example, by a second neural network), allowing the downsampling neural network 204 to access geometric properties of the original texture map 202 to determine parameters for downsampling in a geometry-aware manner. As some examples of the downsampling neural network 204 detecting the plurality of down-sized features 206, the downsampling neural network 204 may perform contour detection, edge detection, facial recognition, or the like to detect the plurality of down-sized features 206. In the example of FIG. 6, the plurality of down-sized features 206 are contours of the original texture map 202.

[0038] The original texture map 202 and the geometric features are processed via an encoderlike module comprising of convolution and pooling layers which extract a set of deep features (Fe) at the desired reduced resolution of H / s x W / s and Ci channels. The reconstruction operation 208 then reconstructs the texture map using the plurality of down-sized features 206 from the downsampling neural network 204 to generate the downsized texture map 210. The plurality of down-sized features 206 within the downsized texture map 210 may be in slightly different locations or may be altered to be associated with more, fewer, and / or different polynomials than within the original texture map 202. In this manner, the operation 200 may result in important features within the original texture map 202 being assigned additional, fewer, and / or different polynomials.

[0039] To improve the parameters of the downsampling neural network 204, the difference between the rendered views generated from the original texture map 202 and the downsized texture map 210 is minimized. For example, a differentiable rendering operation 212 is performed separately on the original texture map 202 and the downsized texture map 210. While typical rendering generates a 2D image from some 3D geometric representations, the differentiable rendering operation 212 pinpoints attributes of the scene while generating the 2D image. In some instances, the differentiable rendering operation 212 identifies geometry elements, illumination parameters, and light fields that affect each pixel in the 2D image.

[0040] At each iteration of training the downsampling neural network 204, a batch PK={Pj] •_1°f camera poses (P7) focusing on the mesh are sampled randomly or pseudo-randomly from a predefined range S of plausible and reasonable viewing conditions. The range S may be received from a user. The range S may be selected such that all reasonable viewing angles and distances are included within it. Additionally, at each iteration of training, K images are rendered from both the original map and the processed map with cameras placed and oriented according to the set of sampled poses PK.P) is the rendering function that projects a textured mesh M(-) to an RGB image from the viewing pose P and I^rigand I^sare the sets of images rendered from the original textured map 202 and the downsized texture map 210, respectively.

[0041] FIG. 3 provides a block diagram of an example differentiable rendering operation 212. Differentiable rendering is performed on scene parameters 300 to generate a rendered image 302. The scene parameters include meshes, textures, lighting, camera positions, and the like. The rendered image 302 is then compared to a reference image 304 to generate a loss image 306. The loss image 306 is indicative of differences between the rendered image 302 and the reference image 304, and more particularly, of losses computed in the rendered image 302. A gradient of losses is computed for each scene parameter 300 and is provided back as feedback to the scene parameters 300. By backpropagating the gradient of losses, the scene attributes for the 3D model associated with the rendered image can be improved.

[0042] In some implementations, for differentiable rendering of the original texture map 202 and the downsized texture map 210, the ambient light is set so that the mesh surface is illuminated evenly. Additionally, the original texture map 202 and the downsized texture map 210 may be diffuse such that colors are independent to viewing direction. Additionally, differentiable rendering of the original texture map 202 and the downsized texture map 210 is performed using the same random set of camera positions. For example, as shown in FIG. 4, four random camera positions may be implemented for differentiable rendering.

[0043] Returning to FIG. 2, the loss image 306 of each differentiable rendering operation 212 is provided to a mean loss function 214. The loss of each pair of rendered images are summed up to obtain the rendering loss for each iteration, defined by Equation 1 :

[0044] In some implementations, the mean loss function 214 identifies the gradient of losses of the loss image 306 of both the original texture map 202 and the downsized texture map 210. The loss is provided to the downsampling neural network 204 as feedback. The performance of the downsampling neural network 204 may be validated every few iterations by averaging the losses computer from images rendered from a set of V number of poses sampled uniformly from PK.

[0045] Providing feedback to the downsampling neural network 204 by the mean loss function 214 may be iterated until a completion condition is satisfied. For example, a predetermined number of iterations may be performed to train the downsampling neural network 204. In other instances, feedback is provided to the downsampling neural network 204 to train the downsampling neural network 204 until the mean loss function 214 begins approaching an asymptote or otherwise shows that additional training will not significantly improve results relative to the additional training costs. In some instances, the training of the downsampling neural network 204 may continue until it is determined that the rate of change of differences between the rendered downsampled texture map and the rendered original texture map over a predetermined number (e.g., the rate of training improvements has dropped below a predetermined rate) of prior iterations is less than a predetermined rate of change or until it is determined that the differences between the rendered downsampled texture map and the rendered original texture map of at least one training iteration is less than a predetermined amount of differences (e.g., after achieving a predetermined level of quality, such as could be measured by structural similarity index measure (SSIM) values, in the downsampled texture relative to the original texture). For example, the training of the downsampling neural network 204 may continue until a SSIM between a rendered downsampled texture map and a rendered texture map exceeds a predetermined structural similarity index measure. As another example, the training of the downsampling neural network 204 may continue until a rate of change of a SSIM between a rendered downsampled texture map and a rendered texture map drops below a predetermined rate of change over a predetermined number of prior generations of training. In these and other examples, the SSIM between a rendered downsampled texture map and a rendered texture map may be measured at multiple (e.g., randomly -selected) camera positions, with SSIMs fromdifferent camera positions being averaged together or otherwise analyzed. In some instances, even when an average SSIM reached a predetermined goal, or the rate of change of SSIM drops below a predetermined goal, training may continue until reaching a minimum predetermined number of generations and / or may be continue if the SSIM at one or more camera positions exceeds a predetermined threshold (i.e., if there is an outlier with a sufficiently worse SSIM).

[0046] In some instances, the downsampling neural network 204 is trained for a particular 3D model prior to outputting a downsampled texture map to module 216 for display and / or distribution. For example, the downsampling neural network 204 is trained for each 3D model for which an associated texture map will be generated. In this manner, the downsampling neural network 204 has an understanding of the geometrical shape of the 3D model. Once the downsampling neural network 204 is trained, the original texture map 202 is provided to the downsampling neural network 204 and the downsized texture map 210 is generated and provided to module 216 for further processing, display, storage, and / or distribution. The downsampling neural network 204 is re-trained for each 3D model and associated original texture map 202.

[0047] Users may view rendered images and 3D models from different viewing directions. Additionally, texture maps rest on non-planar surfaces that may cause discontinuities. Embodiments described herein may further account for geometric information, including geometric nonconformity, while downsampling the original texture map 202. FIG. 5 illustrates an operation 500 of generating texture maps for 3D models that accounts for geometric information. Particularly, in addition to the operations of FIG. 2, the operation 500 includes a geometry module 501 that includes additional processes for altering the plurality of downsized features 206 based on the geometry of the associated 3D model.

[0048] Then downsampling neural network 204 in combination with the geometry module 501 receives the original texture map 202 and accounts for both the geometry layout and the UV parametrization to generate a downsampled texture map which increases the visual fidelity of the rendered views. The geometry module 501 can be expressed mathematically as shown in Equations 2 and 3, where T is the original texture map 202, T 5 is the original texture map 202 downsampled by a factor of s, M is the mesh and r(-) is the rendering function used to transform the 3D mesh into a 2D view. The mesh M includes vertex positions (v), face connectivity (f) and the UV coordinates for all vertices per fact (u).T I s = DownS amplers(T) [Equation 2]T l s = GeoScalersM(T, v, f,u), r(-')) [Equation 3]

[0049] The downsampling neural network 204 is specifically trained to downsample each texture map. Particularly, for every given input texture map, determining the parameters of the downsampling neural network 204 that downsamples the textures of the mesh involves an iterative self-supervised process of minimizing differences between the rendered views of the original texture map 202 and the downsized texture map 210, as determined by mean loss function 214. This process is provided by Equations 4 and 5, where DNN is a function of the downsampling neural network 204:[Equation 4]6opt= argmin0RenderLoss(M(T s, v,f, u''),M(T, v,f,u)') [Equation s]

[0050] The output of the downsampling neural network 204, as previously described, is provided to the geometry module 501 that filters the down-sized features 206 in accordance with the connectivity and the layout of the texture map. Prior to reconstruction of an RGB texture map, a warping function is estimated to be applied to the texture map and the UV parameters of the mesh vertices. The warping increases the efficiency of the texture layout and the UV mapping of the map.

[0051] The geometry module 501 includes a sampling vertex- wise features block 502, a graph convolution block 504, and a re-fill texture map block 506. The sampling vertex-wise features block 502 uses UV parametrization to sample a Ci dimension feature embedded at the location of each of the V mesh vertices from Fe, the Ci x H / s x W / s dimensional feature map produced by the downsampling neural network 204 (e.g., the plurality of down-sized features 206). If a vertex appeal's more than once in the UV parametrization, the samples are averaged from each of its positions. The sampling vertex- wise features block 502 outputs a graph with V vertices and a C dimensional feature embedding associated with each vertex. The vertex wise feature embeddings lie adjacent to their true neighbors on the 3D surface and are connected to their neighbors via the edges of the map. FIG. 7 illustrates an example output 700 of the sampling vertex-wise features block 502.

[0052] The graph convolution block 504 filters the vertex-wise feature embeddings using residual blocks using graph convolutions, such as that provided by Equation 6: ft =9ifi + O2 ' jEN(j)eijfj [Equation 6] where / i is the feature embedding at vertex i, N(i) is the set of neighbors of i, etj is the weight between nodes z and j, and 0i and 02 are learned parameters. The edge weights eij may be set to equal the Euclidean distance between the mesh vertices. By performing the convolutions in the geometric domain, features that are proximal in the 3D domain on the surface of the mesh, but are lying on the distant seams of texture maps, are fused together. FIG. 8 illustrates an example output 800 of the graph convolution block 504.

[0053] The re-fill texture map block 506 reconstructs a 2D map from the geometrically filtered features. The re-fill texture map block 506 first smoothly interpolates the features sampled at each vertex across the faces using the barycentric coordinates. Next, the UV parametrization is used to project the features on faces onto the 2D plane. A 2D feature map (Fg) of size C2 x H / s x W / s, which are compensated for geometric discontinuities and distortions present in the texture map, is the output of the re- fill texture map block 506, and therefore the geometry module 501. The 2D feature map Fgmay henceforth be referred to as a downsampled RGB texture map. In some instances, Fg is added with a skip residual from Fein instances where the vertices are sparse and sampled vectors at the vertices fail to retain sufficient information.

[0054] A reconstruction and UV warping block 508 reconstructs and warps the downsampled RGB texture map from the geometry module 501. Particularly, before reconstructing the downsampled RGB texture map, a warping function fwarp- R2“ * R2isestimated that improves the UV parametrization. This warping improves the layout of the texture map by assigning larger areas to more intricate textures. The warping function fwarpis estimated by deriving a pixel- wise offset for each pixel on the texture map from a slice of the geometrically encoded features Fg. The predicted offset is applied to the remaining slide of Fgto obtain a texture map features Fwwith an improved layout over the original texture map 202. A new UV parametrization u' = fwarp(u) is obtained by applying the estimated warping function on each of the original UV parameters 11. The calculation of new UV parameters is necessary to map the faces to the corresponding regions in newly warped texture map (Fw).

[0055] The UV warping block 508 reconstructs the downsampled RGB texture map T l s from Fw. This downsampled texture map T f s is overlaid on the surface of the mesh using the offset- adjusted UV parameters u’ to obtain the mesh representation M(T f s, v, f w’) as the final output (e.g., downsized texture map 210). FIG. 9A illustrates an example original texture map 202 and FIG. 9B illustrates an example downsized texture map 210. While both the original texture map 202 and downsized texture map 210 include the same features, the features within the downsized texture map 210 may be downsampled at different rates due to the spatially non-uniform nature of the downsampling method described herein. For example, certain features of the original texture map 202 may be identified as being “more important” or having greater detail to preserve during downsampling. These features are assigned a greater relative number of polynomials during generation of the downsized texture map 210, preserving the quality of the feature. In other words, these features may grow in apparent size relative to other features, when the downsized texture map 210 is viewed at the same dimensions as the original texture map 202. Other features may then be assigned relatively fewer polynomials in the downsized texture map 210 compared to the original texture map 202 (e.g., other features may shrink in apparent size relative to other features). In some implementations, the arrangement of polynomials forming the original texture map 202 may change when forming the downsized texture map 210. Additionally, mapping of the polynomials in the downsized texture map 210 to the 3D model may change compared to the original texture map 202, as some features are downsampled more than other features.

[0056] While some areas of interest of indicated within FIG. 9A and 9B for comparison purposes, FIG. 10A provides a detailed view of an original area of interest 900A and FIG. 10B provides a detailed view of a corresponding downsampled area of interest 900B. As shown by comparing the original area of interest 900A and the downsampled area of interest 900B, polynomials forming the texture map may be resized during downsampling. For example, a first polynomial 1000 is increased in size relative to some of its neighboring polynomials in the downsampled area of interest 900B compared to the original area of interest 900A.

[0057] The quality gains provided by operation 500 are quantified in FIGS. 11 A- 1 ID. The operation 500 outperforms existing methods, improving downsampling quality for 4x and 8x scale factors. Additionally, the trade-off cost between memory usage and rendering and visual quality is greatly improved, with higher visual quality of rendered content being available at level lower memory usage. Improvements in visual results are visualized in FIG. 12. In FIGS. 11A-1 ID, PSNRrefers to peak signal-to-noise ratio and SSIM refers to the structural similarity index measure (which quantifies image quality degradation caused by processing such as data compression by comparing the similarity of two images such as an original image and a downsampled image).

[0058] The above video delivery systems and methods may provide for generating textured maps for computer, mobile, and virtual reality graphics applications. Systems, methods, and devices in accordance with the present disclosure may take any one or more of the following configurations.

[0059] (1) A method for generating downsampled texture maps, the method comprising: training a neural network by (i) providing, to the neural network, a three-dimensional (3D) model including an original texture map that covers a 3D mesh, (ii) rendering the original texture map on the 3D mesh and rendering a downsampled texture map on the 3D mesh or on a downsampled version of the 3D mesh based on current parameters of the neural network; (iii) generating loss function feedback by analyzing differences between the rendered downsampled texture map and the rendered original texture map, (iv) using the loss function feedback to alter the parameters of the neural network, and (v) iterating (ii) through (iv) until a completion condition is satisfied; and generating, with the trained neural network, at least one final downsampled texture map, wherein the final downsampled texture map has a lower resolution than the original texture map.

[0060] (2) The method according to (1), wherein the completion condition is a predetermined number of iterations of (ii) through (iv).

[0061] (3) The method according to any one of (1) to (2), wherein the completion condition comprises determining that a structural similarity index measure between a rendered downsampled texture map and a rendered texture map exceeds a predetermined structural similarity index measure.

[0062] (4) The method according to any one of (1) to (2), wherein the completion condition comprises determining that a rate of change of structural similarity index measures between rendered downsampled texture maps and rendered texture maps has dropped below a predetermined rate of change over a predetermined number of prior iterations of (ii) through (iv).

[0063] (5) The method according to any one of (1) to (4), wherein (iii) generating loss function feedback by analyzing differences between the rendered downsampled texture map and the rendered original texture map includes performing a differentiable rendering operation on the renderedoriginal texture map using a first set of camera positions and performing the same differentiable rendering operation on the rendered downsampled texture map using the same first set of camera positions.

[0064] (6) The method according to (5), further comprising pseudo-randomly selecting the first set of camera positions.

[0065] (7) The method according to (5), further comprising receiving, from a user, a set of possible camera positions and pseudo-randomly selecting the first set of camera positions from the received set of possible camera positions.

[0066] (8) The method according to any one of (5) to (7), wherein generating the final downsampled texture map includes sampling a set of the plurality of features located at the vertices of the original texture map using UV parametrization, thereby generating vertex feature embeddings associated with each vertex.

[0067] (9) The method according to (8), wherein generating the final downsampled texture map includes filtering the vertex feature embeddings using graph convolutions, thereby generating a plurality of geometrically filtered features.

[0068] (10) The method according to (9), wherein generating the final downsampled texture map includes reconstructing a two-dimensional image using the plurality of geometrically filtered features.

[0069] (11) The method according to (10), wherein generating the final downsampled texture map includes applying a UV warping function to the two-dimensional image.

[0070] (12) The method according to any one of (1) to (11), wherein generating the final downsampled texture map includes downsampling the original texture map such that some regions of the original texture map are downsized more than others.

[0071] (13) The method according to any one of (1) to (12), wherein the neural network is a convolutional neural network configured to detect contours within the original texture map.

[0072] (14) A system for generating downsampled textures, the system including: an electronic processor connected to a memory, wherein the memory stores a neural network, and wherein theelectronic processor is configured to: train the neural network by (i) providing, to the neural network, a three-dimensional (3D) model including an original texture map that covers a 3D mesh, (ii) rendering the original texture map on the 3D mesh and rendering a downsampled texture map on the 3D mesh or on a downsampled version of the 3D mesh based on current parameters of the neural network; (iii) generating loss function feedback by analyzing differences between the rendered downsampled texture map and the rendered original texture map, (iv) using the loss function feedback to alter the parameters of the neural network, and (v) iterating (ii) through (iv) until a completion condition is satisfied; and generate, with the trained neural network, at least one final downsampled texture map, wherein the final downsampled texture map has a lower resolution than the original texture map.

[0073] (15) The system according to (14), wherein the completion condition is a predetermined number of iterations of (ii) through (iv).

[0074] (16) The system according to any one of (14) to (15), wherein the completion condition comprises determining that a structural similarity index measure between a rendered downsampled texture map and a rendered texture map exceeds a predetermined structural similarity index measure.

[0075] (17) The system according to any one of (14) to (15), wherein the completion condition comprises determining that a rate of change of structural similarity index measures between rendered downsampled texture maps and rendered texture maps has dropped below a predetermined rate of change over a predetermined number of prior iterations of (ii) through (iv).

[0076] (18) The system according to any one of (14) to (17), wherein (iii) generating loss function feedback by analyzing differences between the rendered downsampled texture map and the rendered original texture map includes performing a differentiable rendering operation on the rendered original texture map using a first set of camera positions and performing the same differentiable rendering operation on the rendered downsampled texture map using the same first set of camera positions.

[0077] (19) The system according to (18), wherein the electronic processor is configured to pseudo-randomly select the first set of camera positions.

[0078] (20) The system according to (18), wherein the electronic processor is configured to receive, from a user, a set of possible camera positions and pseudo-randomly select the first set of camera positions from the received set of possible camera positions.

[0079] (21) The system according to any one of claims ( 18)-(20), wherein, to generate the final downsampled texture map, the electronic processor is configured to select a set of the plurality of features located at the vertices of the original texture map using UV parametrization, thereby generating vertex feature embeddings associated with each vertex.

[0080] (22) The system according to (21), wherein, to generate the final downsampled texture map, the electronic processor is configured to filter the vertex feature embeddings using graph convolutions, thereby generating a plurality of geometrically filtered features.

[0081] (23) The system according to (22), wherein, to generate the final downsampled texture map, the electronic processor is configured to reconstruct a two-dimensional image using the plurality of geometrically filtered features.

[0082] (24) The system according to (23), wherein, to generate the final downsampled texture map, the electronic processor is configured to apply a UV warping function to the two-dimensional image.

[0083] (25) The system according to any one of (14) to (24), wherein, to generate the final downsampled texture map, the electronic processor is configured to downsample the original texture map such that some regions of the original texture map are downsized more than others.

[0084] (26) The system according to any one of (14) to (25), wherein the neural network is a convolutional neural network configured to detect contours within the original texture map.

[0085] With regard to the processes, systems, methods, heuristics, etc. described herein, it should be understood that, although the steps of such processes, etc. have been described as occurring according to a certain ordered sequence, such processes could be practiced with the described steps performed in an order other than the order described herein. It further should be understood that certain steps could be performed simultaneously, that other steps could be added, or that certain steps described herein could be omitted. In other words, the descriptions of processesherein are provided for the purpose of illustrating certain embodiments, and should in no way be construed so as to limit the claims.

[0086] Accordingly, it is to be understood that the above description is intended to be illustrative and not restrictive. Many embodiments and applications other than the examples provided would be apparent upon reading the above description. The scope should be determined, not with reference to the above description, but should instead be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. It is anticipated and intended that future developments will occur in the technologies discussed herein, and that the disclosed systems and methods will be incorporated into such future embodiments. In sum, it should be understood that the application is capable of modification and variation.

[0087] All terms used in the claims are intended to be given their broadest reasonable constructions and their ordinary meanings as understood by those knowledgeable in the technologies described herein unless an explicit indication to the contrary is made herein. In particular, use of the singular articles such as “a,” “the,” “said,” etc. should be read to recite one or more of the indicated elements unless a claim recites an explicit limitation to the contrary.

[0088] The Abstract of the Disclosure is provided to allow the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. In addition, in the foregoing Detailed Description, it can be seen that various features are grouped together in various embodiments for the purpose of streamlining the disclosure. This method of disclosure is not to be interpreted as reflecting an intention that the claimed embodiments incorporate more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive subject matter lies in fewer than all features of a single disclosed embodiment. Thus, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as a separately claimed subject matter.

Claims

CLAIMSWhat is claimed is:

1. A method for generating downsampled texture maps, the method comprising: training a neural network by (i) providing, to the neural network, a three-dimensional (3D) model including an original texture map that covers a 3D mesh, (ii) rendering the original texture map on the 3D mesh and rendering a downsampled texture map on the 3D mesh or on a downsampled version of the 3D mesh based on current parameters of the neural network; (iii) generating loss function feedback by analyzing differences between the rendered downsampled texture map and the rendered original texture map, (iv) using the loss function feedback to alter the parameters of the neural network, and (v) iterating (ii) through (iv) until a completion condition is satisfied; and generating, with the trained neural network, at least one final downsampled texture map, wherein the final downsampled texture map has a lower resolution than the original texture map.

2. The method according to claim 1, wherein the completion condition is a predetermined number of iterations of (ii) through (iv).

3. The method according to claim 1 or claim 2, wherein the completion condition comprises determining that a structural similarity index measure between a rendered downsampled texture map and a rendered texture map exceeds a predetermined structural similarity index measure.

4. The method according to claim lor claim 2, wherein the completion condition comprises determining that a rate of change of structural similarity index measures between rendered downsampled texture maps and rendered texture maps has dropped below a predetermined rate of change over a predetermined number of prior iterations of (ii) through (iv).

5. The method according to any one of claims 1-4, wherein (iii) generating loss function feedback by analyzing differences between the rendered downsampled texture map and the rendered original texture map includes performing a differentiable rendering operation on the rendered original texture map using a first set of camera positions and performing the same differentiablerendering operation on the rendered downsampled texture map using the same first set of camera positions.

6. The method according to claim 5, further comprising pseudo-randomly selecting the first set of camera positions.

7. The method according to claim 5, further comprising receiving, from a user, a set of possible camera positions and pseudo-randomly selecting the first set of camera positions from the received set of possible camera positions.

8. The method according to any one of claims 5-7, wherein generating the final downsampled texture map includes sampling a set of the plurality of features located at the vertices of the original texture map using UV parametrization, thereby generating vertex feature embeddings associated with each vertex.

9. The method according to claim 8, wherein generating the final downsampled texture map includes filtering the vertex feature embeddings using graph convolutions, thereby generating a plurality of geometrically filtered features.

10. The method according to claim 9, wherein generating the final downsampled texture map includes reconstructing a two-dimensional image using the plurality of geometrically filtered features.

11. The method according to claim 10, wherein generating the final downsampled texture map includes applying a UV warping function to the two-dimensional image.

12. The method according to any one of claims 1-11, wherein generating the final downsampled texture map includes downsampling the original texture map such that some regions of the original texture map are downsized more than others.

13. The method according to any one of claims 1-12, wherein the neural network is a convolutional neural network configured to detect contours within the original texture map.

14. A system for generating downsampled textures, the system including: an electronic processor connected to a memory, wherein the memory stores a neural network, and wherein the electronic processor is configured to: train the neural network by (i) providing, to the neural network, a three-dimensional (3D) model including an original texture map that covers a 3D mesh, (ii) rendering the original texture map on the 3D mesh and rendering a downsampled texture map on the 3D mesh or on a downsampled version of the 3D mesh based on current parameters of the neural network; (iii) generating loss function feedback by analyzing differences between the rendered downsampled texture map and the rendered original texture map, (iv) using the loss function feedback to alter the parameters of the neural network, and (v) iterating (ii) through (iv) until a completion condition is satisfied; and generate, with the trained neural network, at least one final downsampled texture map, wherein the final downsampled texture map has a lower resolution than the original texture map.

15. The system according to claim 14, wherein the completion condition is a predetermined number of iterations of (ii) through (iv).

16. The system according to claim 14 or claim 15, wherein the completion condition comprises determining that a structural similarity index measure between a rendered downsampled texture map and a rendered texture map exceeds a predetermined structural similarity index measure.

17. The system according to claim 14 or claim 15, wherein the completion condition comprises determining that a rate of change of structural similarity index measures between rendered downsampled texture maps and rendered texture maps has dropped below a predetermined rate of change over a predetermined number of prior iterations of (ii) through (iv).

18. The system according to any one of claims 14-17, wherein (iii) generating loss function feedback by analyzing differences between the rendered downsampled texture map and the rendered original texture map includes performing a differentiable rendering operation on the rendered original texture map using a first set of camera positions and performing the same differentiable rendering operation on the rendered downsampled texture map using the same first set of camera positions.

19. The system according to claim 18, wherein the electronic processor is configured to pseudo- randomly select the first set of camera positions.

20. The system according to claim 18, wherein the electronic processor is configured to receive, from a user, a set of possible camera positions and pseudo-randomly select the first set of camera positions from the received set of possible camera positions.

21. The system according to any one of claims 18 to 20, wherein, to generate the final downsampled texture map, the electronic processor is configured to select a set of the plurality of features located at the vertices of the original texture map using UV parametrization, thereby generating vertex feature embeddings associated with each vertex.

22. The system according to claim 21, wherein, to generate the final downsampled texture map, the electronic processor is configured to filter the vertex feature embeddings using graph convolutions, thereby generating a plurality of geometrically filtered features.

23. The system according to claim 22, wherein, to generate the final downsampled texture map, the electronic processor is configured to reconstruct a two-dimensional image using the plurality of geometrically filtered features.

24. The system according to claim 23, wherein, to generate the final downsampled texture map, the electronic processor is configured to apply a UV warping function to the two-dimensional image.

25. The system according to any one of claims 14 to 24, wherein, to generate the final downsampled texture map, the electronic processor is configured to downsample the original texture map such that some regions of the original texture map are downsized more than others.

26. The system according to any one of claims 14 to 25, wherein the neural network is a convolutional neural network configured to detect contours within the original texture map.

Citation Information

Patent Citations

  • Three-dimensional model attribute image simplification method and device based on rate distortion optimization

    CN114742943A

  • Texture image generation method and device, equipment, storage medium and product

    CN115830091A