A Method and System for Improving Texture Resolution Consistency in 3D Models Based on Arbitrary Scale Super-Resolution

By using an arbitrary-scale super-resolution network enhanced with multi-view information, the super-resolution factor of triangular facet clusters is adaptively calculated and multi-view image information is integrated, which solves the problem of inconsistent resolution of 3D texture models and improves the visual effect of 3D models.

CN119904564BActive Publication Date: 2025-10-31WUHAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411871852.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-10-31
Estimated Expiration
2044-12-18

AI Technical Summary

Technical Problem

Existing technologies struggle to address the issue of inconsistent texture resolution on 3D texture models, which negatively impacts visual quality.

Method used

An arbitrary-scale super-resolution network with multi-view information enhancement is adopted. The super-resolution factor of triangular facet clusters is adaptively calculated, and a cross-view feature aggregation module is introduced to integrate multi-view image information to improve texture resolution consistency.

Benefits of technology

This achieves a more consistent resolution among adjacent triangular facets on the 3D model, thus improving the visualization effect of the 3D model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119904564B_ABST
    Figure CN119904564B_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for improving the texture resolution consistency of 3D models based on arbitrary-scale super-resolution. The method includes: dividing a mesh-formatted 3D geometric module into multiple triangular facet clusters and generating an initial texture atlas; introducing a scale factor to represent the texture resolution of the triangular facet clusters, and then calculating the super-resolution factor of the triangular facet clusters based on the scale factor; adaptively improving the resolution of the texture patches corresponding to each triangular facet cluster using an arbitrary-scale super-resolution network based on reference information, according to the super-resolution factor; after each texture patch has undergone arbitrary-scale super-resolution processing, packaging the super-resolution texture patches to generate a new texture atlas, based on which the final 3D texture model can be generated. This invention not only improves texture resolution but also ensures the consistency of texture resolution between adjacent triangular facet clusters, resulting in a 3D model with consistent texture resolution and realistic visual effects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of 3D real scene model reconstruction technology, and in particular relates to a method and system for improving the texture resolution consistency of 3D models based on arbitrary scale super-resolution. Background Technology

[0002] Realistic 3D models are widely used in smart cities, virtual / augmented reality, urban planning, and other fields. Image-based 3D reconstruction has attracted widespread attention from researchers due to the convenience and low cost of image data acquisition. The image-based 3D reconstruction process can be broadly divided into geometric reconstruction and texture mapping. Geometric reconstruction typically employs techniques such as sparse matching and dense matching to reconstruct the 3D surface model of the target from multiple view images. This model is usually called a "white model"—it only has geometric shapes and no texture information. Texture mapping assigns color and texture to the "white model" based on camera intrinsic and extrinsic parameters and multiple view images, giving the 3D model rich details and a more realistic feel. Specifically, the "white model" is composed of a series of triangular faces, and texture mapping requires assigning texture information to each triangular face based on multi-view image information. According to the method of generating triangular face textures, texture mapping methods can be divided into fusion-based methods and tessellation-based methods. Fusion-based methods perform weighted fusion of all images corresponding to the triangular face to generate the color of the triangular face; tessellation-based methods select the best view image for the triangular face from all candidate images. Compared to fusion-based methods, tessellation-based methods exhibit less deformation when resampling multi-view images into texture space, resulting in textures with better realism and finer details. To improve the local consistency of 3D texture models, tessellation-based methods typically merge adjacent triangles with the same best-view image into triangle clusters, using these clusters as basic units to extract texture tile information from multi-view images.

[0003] When the shooting angle, distance, and camera parameters of multi-view images differ, the texture resolution of adjacent triangular facets on a 3D texture model, which acquire texture information from different multi-view images, exhibits significant differences. This problem is common and greatly degrades the visual quality of the texture model. Although many traditional texture mapping methods have been proposed in recent years to generate texture models, most of these methods focus on camera pose optimization, viewpoint selection, and color correction. While some works have focused on improving the texture resolution of 3D models, they can improve resolution and recover detailed information through fixed-scale super-resolution of texture atlases, but the problem of resolution inconsistency still exists on the texture model. Summary of the Invention

[0004] To address the issue of inconsistent sharpness in 3D texture models, existing technologies cannot effectively resolve this problem. This invention innovatively proposes an arbitrary-scale super-resolution network with multi-view information enhancement. This network integrates multi-view image information to improve the resolution consistency of the texture model, ultimately enhancing the resolution of adjacent triangular facets on the 3D model to the same level, thereby improving the visualization effect of the 3D model. While many arbitrary-scale image super-resolution methods have been proposed, these methods are designed for 2D images rather than 3D texture models. In this invention, the mesh-formatted 3D model is divided into several triangular facets, and the texture patch corresponding to each triangular facet is used as the processing primitive, rather than the entire texture atlas. Since the resolutions of different triangular facets are inconsistent and arbitrary in scale, the scale factor of each triangular facet is adaptively calculated, which determines the final resolution of the corresponding texture patch. That is, different triangular facets have different super-resolution factors. Furthermore, considering the ill-posedness of single-image super-resolution and the requirement of multi-view consistency in the super-resolution results, a cross-view feature aggregation module is designed to integrate multi-view image information to improve the performance of the super-resolution network. Thanks to the adaptive calculation of super-resolution factors and the fusion of multi-view image information, the texture resolution of the 3D model can be improved and the resolution of adjacent triangular facets can be made more consistent.

[0005] The technical solution of this invention is: a method for improving the texture resolution consistency of a 3D model based on arbitrary scale super-resolution, comprising the following steps:

[0006] Step 1: Determine the optimal viewpoint image corresponding to each triangle face, and generate a triangle face cluster and an initial texture atlas composed of texture tiles;

[0007] Step 2: Calculate the super-resolution factor of the triangular facet cluster based on the area of ​​the triangular facet cluster and the texture tile.

[0008] Step 3: Use a super-resolution network to improve the resolution of the texture tiles corresponding to the triangular facet clusters;

[0009] The super-resolution object is the smallest bounding rectangle region of the texture tile in its optimal view image. The super-resolution network generates super-resolution tiles based on the super-resolution factor. Reference image information is introduced to achieve higher quality super-resolution reconstruction. The reference image is a multi-view image that is spatially adjacent to the optimal view image.

[0010] Step 4: After each texture tile has undergone arbitrary scale super-resolution processing, the super-resolution texture tiles are packaged to generate a new texture atlas. Based on the updated texture atlas, a 3D model with consistent texture resolution and realistic visual effects is generated.

[0011] Furthermore, in step 1, to generate the triangular facet cluster, we first determine the value of each triangular facet f. i The corresponding optimal viewing angle image is denoted as image label l. i Specifically, an energy equation E(l) is constructed to select the optimal viewing angle image, where the data term E data The smoothing term E is used to select a better visible viewpoint image for the triangular facet. smooth Used to minimize the visibility of visual seams; the energy equation E(l) is expressed as:

[0012] E(l)=E data +E smooth

[0013]

[0014] Where l represents the set of image labels for all triangular faces, Faces represents the set of all triangular faces, Edge represents the edge formed by two adjacent triangular faces, and Edges represents the set of all edges. j Describing the triangular face f j The optimal viewing angle image, j is used to distinguish it from i, j≠i, φ(f i ,l i ) represents the triangular face f i In images The projection on Let [·] denote the Sobel operator and [·] denote the Iverson bracket. Solve the above energy equation using the graph cut method to obtain the labels of the optimal viewpoint images of all triangular faces. Aggregate adjacent triangular faces with the same label into a triangular face cluster. Where C represents the number of triangular facets; subsequently, texture tiles corresponding to the triangular facets are obtained from the multi-view image, packaged to generate a texture atlas, and all texture tiles are denoted as...

[0015] Furthermore, the specific implementation method of step 2 is as follows:

[0016] Step 2.1, Generation of the scale factor for the triangular facet family: scale factor r i Defined as a triangular cluster c i The area and the texture patch t corresponding to the triangular cluster i The ratio of their areas is shown below:

[0017]

[0018] in and They represent c respectively i and t i The area ∈ is a very small non-zero constant; r iThe larger the value, the clearer the corresponding texture on the 3D model;

[0019] Step 2.2, Super-resolution factor generation for triangular facets: Since the texture resolution of different triangular facets varies, the super-resolution factor of each triangular facet is adaptively calculated based on the scale factor defined in equation (2); before calculating the super-resolution factor, a reference scale factor r is set. ref Then, the super-resolution factor s of each triangular cluster. i We obtain it from the following formula:

[0020]

[0021] Furthermore, the reference scaling factor r ref r is the maximum value of the scale factor for all triangular facets. max Or greater than r max The value of .

[0022] Furthermore, the specific implementation method of step 3 is as follows:

[0023] Step 3.1, Feature Extraction: The super-resolution network EDSR-b is used as a feature extractor to generate texture patches t. LR Feature map F LR and reference image feature map J represents the number of reference images;

[0024] Among them, t LR For texture tile t i Image from its optimal viewing angle The smallest enclosing rectangular region in;

[0025] Step 3.2, Cross-view feature aggregation: For feature map F LR For each pixel p in the image, it is first projected onto the reference image viewpoint according to multi-view geometry. The projected point is denoted as p. Also known as a query point, Create a small window in the center A cross-attention mechanism is used to enhance the image features of point p, generating the final enhanced feature map. To integrate reference information from multi-view images into low-resolution image feature maps;

[0026] Step 3.3, Intra-view information integration: The arbitrary-scale super-resolution method CiaoSR is used to integrate intra-view information to obtain the implicit encoding corresponding to the coordinates of the query point;

[0027] Step 3.4, Super-resolution image reconstruction: For each pixel coordinate in the super-resolution result image, its color value is predicted by combining the coordinates of the query point and its implicit encoding through an implicit image function. Then, the color prediction values ​​of all pixel coordinates are added pixel by pixel to the bicubic interpolation result image of the low-resolution image to generate the super-resolution image. Subsequently, the effective texture region in the super-resolution image is obtained through masking, and the texture patch after resolution improvement is obtained.

[0028] Step 3.5, using the loss function: train the super-resolution network at any scale based on the training sample set until the entire network converges to the optimal accuracy.

[0029] Furthermore, in step 3.2, a cross-attention mechanism is used to enhance the image features of point p, and the calculation formula is shown below:

[0030]

[0031] in, F represents the feature of point p after the reference information is aggregated through the cross-attention mechanism. inter For the corresponding feature map; (m,n) indicates that it is contained in the window. The coordinates within the range, σ represents the Softmax function; Q, and Let represent the query, key, and value in the attention mechanism, respectively, and their definitions are as follows:

[0032]

[0033] in, For low-resolution image features of pixel p, For the reference image features at coordinates (m,n), dis m,n For projection point The relative distance between the cell and its neighboring pixels (m,n) within the window; cell is the grid decoded to specify the shape of the query region; φ k and φ v The key network and value network are implemented by a multilayer perceptron, both of which are used for feature mapping; according to equation (4), the features of all reference images of pixel p are aggregated. To generate the final enhanced feature map Further integration of F inter and F LR First, the two feature maps are concatenated along the channel dimension. Then, they are processed by two convolutional layers (Conv) and two non-linear activation functions (ReLU). The specific process of the feature fusion module is as follows:

[0034]

[0035] Furthermore, the specific implementation method of step 3.3 is as follows;

[0036] Assume x is the super-resolution result image t SR pixel coordinates, x q For low-resolution image t LR The corresponding query point coordinates in x q small local area centered on Execute in-view attention mechanism to obtain implicit encoding Formulated as:

[0037]

[0038] The query, key, and value in the in-view attention mechanism are defined as follows:

[0039]

[0040] In the formula (m * ,n * (x) is the closest query point. q pixels, and Indicates in (m * ,n * The enhanced low-resolution features at positions (m,n) and (m,n), φ k ′ and φ v ′ represents the key network and the value network, respectively. This represents the enhanced low-resolution feature map. Non-local features captured in the process.

[0041] Furthermore, in step 3.4, the super-resolution image t SR The calculation is shown in the following formula:

[0042]

[0043] in, The graph shows the result of bicubic interpolation, f θ (z x (x) is an implicit image function, implemented using a multilayer perceptron, z x This represents the implicit encoding corresponding to the x-coordinate of each pixel in the super-resolution image.

[0044] Furthermore, in step 3.5, the loss function is... The norm constrains the differences between the super-resolution image and the high-resolution image at the pixel level.

[0045] The present invention also provides a system for improving the texture resolution consistency of a 3D model based on arbitrary scale super-resolution, including a processor and a memory. The memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute the method for improving the texture resolution consistency of a 3D model based on arbitrary scale super-resolution as described in the above technical solution.

[0046] Compared with the prior art, the present invention has the following three advantages:

[0047] 1) The problem of improving the resolution consistency of 3D texture models is transformed into a texture super-resolution task of arbitrary scale, and a method is proposed to solve it.

[0048] 2) Introduce a super-resolution factor to achieve adaptive improvement of the resolution of triangular facet clusters, so as to improve the resolution consistency of the texture model while improving the texture resolution.

[0049] 3) A cross-view feature aggregation module was designed in the arbitrary scale super-resolution network to better integrate multi-view image information, thereby improving the fidelity of the super-resolution result map and the consistency of content between viewpoints. Attached Figure Description

[0050] Figure 1 This is a diagram of the framework for improving the texture resolution consistency of 3D models based on arbitrary scale super-resolution proposed in this invention.

[0051] Figure 2 This invention provides an arbitrary-scale super-resolution network for enhancing multi-view image information.

[0052] Figure 3 This is a schematic diagram of the experimental results of the arbitrary scale super-resolution network in this invention on test data. From left to right, the effects of different super-resolution scales are shown.

[0053] Figure 4 The diagram shows the experimental results of the 3D model texture resolution consistency improvement method based on arbitrary scale super-resolution in this invention on the test dataset. From left to right, the results are: texture mapping using the original low-resolution texture atlas, texture mapping using the default reference scale factor after super-resolution of the texture atlas, and texture mapping using a larger scale factor after super-resolution of the texture atlas. Detailed Implementation

[0054] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the embodiments described herein are for illustrative and explanatory purposes only and are not intended to limit the scope of the invention.

[0055] This invention provides a method such as Figure 1The method shown is a texture resolution consistency improvement method for 3D models based on arbitrary scale super-resolution, given a 3D geometric model M (with vertices). and triangles A series of multi-view images Based on its camera parameters (V, F, N represent the number of vertices, the number of triangles, and the number of multi-view images, respectively), this invention aims to generate a 3D model with consistent texture resolution. When generating this 3D model, it is necessary to address the problem of inconsistent texture resolution caused by inconsistent multi-view image resolution, thereby generating a 3D model with better visual effects. To achieve this goal, this invention proposes a method for improving the consistency of 3D model texture resolution based on arbitrary-scale super-resolution. The method proposed in this invention can adaptively improve the texture resolution of each triangle cluster, eliminate visual seams caused by inconsistent texture resolution, and improve the visual effect of the 3D model. The technical solution adopted in this invention includes the following steps:

[0056] Step 1: Generation of Triangle Facet Clusters and Texture Atlas. This invention employs a tessellation-based texture mapping method, MVS-Texturing, to generate triangle facet clusters and an initial texture atlas. To generate the triangle facet clusters, it is first necessary to determine the value of each triangle facet f. i The corresponding optimal viewing angle image is denoted as image label l. i Specifically, an energy equation E(k) is constructed to select the optimal viewing angle image, where the data term E... data The smoothing term E is used to select a better visible viewpoint image for the triangular facet. smooth This is used to minimize the visibility of visual seams. The energy equation E(l) is expressed as:

[0057]

[0058] Where l represents the set of image labels for all triangular faces, Faces represents the set of all triangular faces, Edge represents the edge formed by two adjacent triangular faces, and Edges represents the set of all edges. j Describing the triangular face f j The optimal viewing angle image, j is used to distinguish it from i, j≠i, φ(f i ,l i ) represents the triangular face f i In images The projection on Let represent the Sobel operator, and [·] represent Iverson brackets. Solving the energy equation using the graph cut method yields the labels for the optimal viewpoint images of all triangular faces. Adjacent triangular faces with the same label are then grouped into triangular face clusters. Where C represents the number of triangular facets. Subsequently, texture tiles corresponding to the triangular facets are extracted from the multi-view image and packaged to generate a texture atlas. All texture tiles are denoted as...

[0059] Step 2: Calculation of the super-resolution factor for triangular facet clusters. To standardize the resolution of triangular facet clusters, this invention introduces a scale factor to represent the texture resolution of the triangular facet clusters. The super-resolution factor of the triangular facet clusters is then calculated based on the scale factor. Specifically,

[0060] Step 2.1, generating the scale factor for the triangular facet cluster. Scale factor r i Defined as a triangular cluster c i The area and the texture patch t corresponding to the triangular cluster i The ratio of their areas is shown below:

[0061]

[0062] in and They represent c respectively i and t i The area of ​​∈ is a very small constant. i The larger the value, the clearer the corresponding texture on the 3D model.

[0063] Step 2.2, Super-resolution factor generation for triangular facet clusters. Since different triangular facet clusters have different texture resolutions, this invention adaptively calculates the super-resolution factor for each triangular facet cluster based on the scale factor defined in equation (2). Before calculating the super-resolution factor, a reference scale factor r needs to be set. ref In this invention, r ref The default value is the maximum value of the scale factor r of all triangular facet clusters. max .Right now Then, the super-resolution factor s of each triangular cluster i It can then be obtained through the following formula:

[0064]

[0065] It is worth noting that the reference scaling factor r ref It can also be set to be greater than r max The value of this will further improve the overall resolution of the texture model.

[0066] Step 3: Arbitrary-scale super-resolution based on reference information. After obtaining the super-resolution factor for each triangular facet cluster, a super-resolution network is needed to improve the resolution of the texture tiles corresponding to the triangular facet clusters. Considering that an arbitrary-scale super-resolution network can handle arbitrary scale factors with a single model, this invention uses an arbitrary-scale super-resolution network to improve the resolution of texture tiles. Furthermore, the effective region of a texture tile is usually irregular. To better utilize two-dimensional convolution and integrate neighborhood information to achieve super-resolution, the object to be super-resolution in this invention is the texture tile t. i Image from its optimal viewing angle The smallest enclosing rectangular region in t is called t i For simplicity, in a super-resolution network, patch t... ′ ′ is called t LR Super-resolution networks will be based on the super-resolution factor s. i To generate super-resolution tiles t SR Considering the ill-posedness of single-image super-resolution and the requirement of multi-view consistency for the super-resolution results in this invention, reference image information is introduced to achieve higher-quality super-resolution reconstruction. Reference Image It is with the optimal viewing angle image In spatially proximate multi-view images, J represents the number of reference images. The multi-view geometric relationships used to integrate reference information are known in the texture mapping task, enabling rapid matching of corresponding points between the optimal view image and the reference images.

[0067] The arbitrary-scale super-resolution network based on reference information proposed in this invention is as follows: Figure 2 As shown, it includes the following sub-steps:

[0068] Step 3.1, Feature Extraction. The classic and lightweight super-resolution network EDSR-b is used as the feature extractor for the images to generate local patches t. LR Feature map F LR and reference image feature map

[0069] Step 3.2, cross-view feature aggregation. For feature map F LR For each pixel p in the image, it is first projected onto the reference image viewpoint according to multi-view geometry. The projected point is denoted as p. Also known as the query point. Because camera pose and geometric models may be inaccurate, There is a deviation between the actual projected position of p in the reference image view and the actual projected position of p. Create a small window in the center This involves aggregating information from the reference image. Specifically, a cross-attention mechanism is used to enhance the image features at point p, as shown below:

[0070]

[0071] in, F represents the feature of point p after the reference information is aggregated through the cross-attention mechanism. inter This represents the corresponding feature map. (m,n) indicates that it is contained within the window. The coordinates within the range, where σ represents the Softmax function. Q, and Let represent the query, key, and value in the attention mechanism, respectively, and their definitions are as follows:

[0072]

[0073] in, For low-resolution image features of pixel p, The reference image features are located at coordinates (m,n). It is based on pixel p from the low-resolution feature map F LR Obtained Based on coordinates (m, n), from the reference image feature map Obtained, dis m,n For projection point The relative distance between it and its neighboring pixels (m,n) within the window, i.e. `cell` is the grid decoding used to specify the shape of the query region; `cell` is based on the super-resolution factor `s`. i The calculation yields a specific value of [2 / s]. i ,2 / s i ]. φ k and φ v The key network and value network are implemented using a multilayer perceptron, both used for feature mapping. According to equation (4), the features of all reference images of pixel p are aggregated. To generate the final enhanced feature map Further integration of F inter and F LR First, the two feature maps are concatenated along the channel dimension, and then processed using two convolutional layers (Conv) and two non-linear activation functions (ReLU). The specific process of the feature fusion module is as follows:

[0074]

[0075] The above method can be used to integrate reference information from multi-view images into the feature map of low-resolution images.

[0076] Step 3.3, Intra-view Information Integration. This invention employs the arbitrary-scale super-resolution method CiaoSR to integrate intra-view information. CiaoSR is a super-resolution method based on implicit neural representations. Such methods integrate information using an implicit image function f based on the coordinates of the query point and its implicit encoding. θ (Implemented by a multilayer perceptron) The color value of the coordinate point is predicted. To obtain information-rich implicit encoding, CiaoSR employs an attention mechanism when integrating neighborhood information of the query point, jointly considering the spatial distance and feature similarity between coordinate points. Specifically, assuming x is the super-resolution result image t SR pixel coordinates, x q For low-resolution image t LR The coordinates of the corresponding query point. In x... q small local area centered on Execute in-view attention mechanism to obtain implicit encoding It can be formalized as follows:

[0077]

[0078] The query, key, and value in the in-view attention mechanism are defined as follows:

[0079]

[0080] In the formula (m * ,n * (x) is the closest query point. q pixels, and Indicates in (m * ,n * The enhanced low-resolution features at positions (m,n) are respectively based on the coordinates (m) * ,n * ) and coordinates (m,n) from the enhanced low-resolution feature map Obtained, φ k ′ and φ v ′ represents the key network and the value network, respectively. This represents the enhanced low-resolution feature map. The non-local features captured in the image are generated using a non-local module.

[0081] Step 3.4, Super-resolution Image Reconstruction. For each pixel coordinate x in the super-resolution image, its implicit encoding is obtained. Then, through the implicit image function f θ (z xThe color value is predicted using (x, y) (an implicit image function is typically implemented using a multilayer perceptron). Then, the predicted color values ​​for all pixel coordinates are added pixel-by-pixel to the bicubic interpolation result of the low-resolution image, as shown in Figure t. LR↑ In this process, a super-resolution image t is generated. sR As shown in the following formula:

[0082]

[0083] Subsequently, the effective texture regions in the super-resolution image are obtained through masking, thus yielding texture tiles with improved resolution.

[0084] Step 3.5, using the loss function. Train the super-resolution network at any scale based on the training sample set. The loss function is: The norm constrains the pixel-level difference between the super-resolution image and the high-resolution image. The proposed arbitrary-scale super-resolution network is trained using a loss function until the entire network converges to optimal accuracy.

[0085] Step 4: Texture atlas update and 3D texture model generation. After each texture tile has undergone arbitrary-scale super-resolution processing, the super-resolution texture tiles are... The system packages and generates a new texture atlas. Based on the updated texture atlas, it can generate visually realistic 3D models with consistent texture resolution.

[0086] This invention conducted experiments on image super-resolution at arbitrary scales and texture resolution consistency improvement on test data from the Multi-View Geometry Dataset (DTU). Examples of the results are shown below. Figure 3 and Figure 4 As shown. Figure 3 The images show the results of low-resolution images at different super-resolution scales. From these images, we can see that the present invention can stably improve the resolution of images, and can correctly reproduce the details in low-resolution images even in super-resolution results at a larger scale of 4.0×. Figure 4 The images show the results of texture mapping on the same 3D geometric model using texture atlases of different resolutions. Through three close-up views of the 3D model, we can see that this invention improves the model's resolution by using a default reference scale factor to super-resolution the texture atlas before texture mapping, effectively alleviating the resolution inconsistency problem present in texture mapping results obtained using the original low-resolution texture atlas. Furthermore, this invention uses a scale factor larger than the default reference scale factor to super-resolution the texture atlas before texture mapping; this results in an even greater improvement in the 3D model's resolution while maintaining good model resolution consistency.

[0087] On the other hand, embodiments of the present invention also provide a three-dimensional model texture resolution consistency improvement system based on arbitrary scale super-resolution, including a processor and a memory. The memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute the three-dimensional model texture resolution consistency improvement method based on arbitrary scale super-resolution as described in the above technical solution.

[0088] It should be understood that the above description of the preferred embodiments is quite detailed, but it should not be considered as a limitation on the scope of protection of this invention. Those skilled in the art, under the guidance of this invention, can make substitutions or modifications without departing from the scope of protection of the claims of this invention, and all such substitutions or modifications fall within the scope of protection of this invention. The scope of protection of this invention should be determined by the appended claims.

Claims

1. A method for improving the texture resolution consistency of 3D models based on arbitrary-scale super-resolution, characterized in that: Includes the following steps: Step 1: Determine the optimal viewpoint image corresponding to each triangle face, and generate a triangle face cluster and an initial texture atlas composed of texture tiles; Step 2: Calculate the super-resolution factor of the triangular facet cluster based on the area of ​​the triangular facet cluster and the texture tile. Step 3: Use a super-resolution network to improve the resolution of the texture tiles corresponding to the triangular facet clusters; The super-resolution object is the smallest bounding rectangle region of the texture tile in its optimal view image. The super-resolution network generates super-resolution tiles based on the super-resolution factor. Reference image information is introduced to achieve higher quality super-resolution reconstruction. The reference image is a multi-view image that is spatially adjacent to the optimal view image. Step 4: After each texture tile has undergone arbitrary scale super-resolution processing, the super-resolution texture tiles are packaged to generate a new texture atlas. Based on the updated texture atlas, a 3D model with consistent texture resolution and realistic visual effects is generated.

2. The method for improving the texture resolution consistency of a 3D model based on arbitrary-scale super-resolution as described in claim 1, characterized in that: In step 1, to generate the triangular facet family, we first determine the value of each triangular facet f. i The corresponding optimal viewing angle image is denoted as image label l. i Specifically, an energy equation E(l) is constructed to select the optimal viewing angle image, where the data term E data The smoothing term E is used to select a better visible viewpoint image for the triangular facet. smooth Used to minimize the visibility of visual seams; the energy equation E(l) is expressed as: E(l)=E data +E smooth Where l represents the set of image labels for all triangular faces, Faces represents the set of all triangular faces, Edge represents the edge formed by two adjacent triangular faces, and Edges represents the set of all edges. j Describing the triangular face f j The optimal viewing angle image, j is used to distinguish it from i, j≠i, φ(f i , l i ) represents the triangular face f i In images The projection on Let [·] denote the Sobel operator and [·] denote the Iverson bracket. Solve the above energy equation using the graph cut method to obtain the labels of the optimal viewpoint images of all triangular faces. Aggregate adjacent triangular faces with the same label into a triangular face cluster. Where C represents the number of triangular facets; subsequently, texture tiles corresponding to the triangular facets are obtained from the multi-view image, packaged to generate a texture atlas, and all texture tiles are denoted as...

3. The method for improving the texture resolution consistency of a 3D model based on arbitrary-scale super-resolution as described in claim 1, characterized in that: The specific implementation method of step 2 is as follows: Step 2.1, Generation of the scale factor for the triangular facet family: scale factor r i Defined as a triangular cluster c i The area and the texture patch t corresponding to the triangular cluster i The ratio of their areas is shown below: in and They represent c respectively i and t i The area ∈ is a constant; r i The larger the value, the clearer the corresponding texture on the 3D model; Step 2.2, Super-resolution factor generation for triangular facets: Since the texture resolution of different triangular facets varies, the super-resolution factor of each triangular facet is adaptively calculated based on the scale factor defined in equation (2); before calculating the super-resolution factor, a reference scale factor r is set. ref Then, the super-resolution factor s of each triangular cluster. i We obtain it from the following formula:

4. The method for improving the texture resolution consistency of a 3D model based on arbitrary-scale super-resolution as described in claim 3, characterized in that: Reference scale factor r ref r is the maximum value of the scale factor for all triangular facets. max Or greater than r max The value of .

5. The method for improving the texture resolution consistency of a 3D model based on arbitrary-scale super-resolution as described in claim 1, characterized in that: The specific implementation method of step 3 is as follows: Step 3.1, Feature Extraction: The super-resolution network EDSR-b is used as a feature extractor to generate texture patches t. LR Feature map F LR and reference image feature map J represents the number of reference images; Among them, t LR For texture tile t i Image from its optimal viewing angle The smallest enclosing rectangular region in; Step 3.2, Cross-view feature aggregation: For feature map F LR For each pixel p in the image, it is first projected onto the reference image viewpoint according to multi-view geometry. The projected point is denoted as p. Also known as a query point, Create a small window in the center A cross-attention mechanism is used to enhance the image features of point p, generating the final enhanced feature map. To integrate reference information from multi-view images into low-resolution image feature maps; Step 3.3, Intra-view information integration: The arbitrary-scale super-resolution method CiaoSR is used to integrate intra-view information to obtain the implicit encoding corresponding to the coordinates of the query point; Step 3.4, Super-resolution image reconstruction: For each pixel coordinate in the super-resolution result image, its color value is predicted by combining the coordinates of the query point and its implicit encoding through an implicit image function. Then, the color prediction values ​​of all pixel coordinates are added pixel by pixel to the bicubic interpolation result image of the low-resolution image to generate the super-resolution image. Subsequently, the effective texture region in the super-resolution image is obtained through masking, and the texture patch after resolution improvement is obtained. Step 3.5, using the loss function: train the super-resolution network at any scale based on the training sample set until the entire network converges to the optimal accuracy.

6. The method for improving the texture resolution consistency of a 3D model based on arbitrary-scale super-resolution as described in claim 5, characterized in that: Step 3.2 uses a cross-attention mechanism to enhance the image features of point p, and the calculation formula is shown below: in, F represents the feature of point p after the reference information is aggregated through the cross-attention mechanism. inter For the corresponding feature map; (m, n) indicates that it is contained in the window. The coordinates within the range, σ represents the Softmax function; Q, and Let represent the query, key, and value in the attention mechanism, respectively, and their definitions are as follows: in, For low-resolution image features of pixel p, For the reference image features at coordinates (m, n), dis m,n For projection point The relative distance between the cell and its neighboring pixels (m, n) within the window; cell is the grid decoded to specify the shape of the query region; φ k and φ v The key network and value network are implemented by a multilayer perceptron, both of which are used for feature mapping; according to equation (4), the features of all reference images of pixel p are aggregated. To generate the final enhanced feature map Further integration of F inter and F LR First, the two feature maps are concatenated along the channel dimension. Then, they are processed by two convolutional layers (Conv) and two non-linear activation functions (ReLU). The specific process of the feature fusion module is as follows:

7. The method for improving the texture resolution consistency of a 3D model based on arbitrary-scale super-resolution as described in claim 5, characterized in that: The specific implementation method of step 3.3 is as follows; Assume x is the super-resolution result image t SR pixel coordinates, x q For low-resolution image t LR The corresponding query point coordinates in x q small local area centered on Execute in-view attention mechanism to obtain implicit encoding Formulated as: The query, key, and value in the in-view attention mechanism are defined as follows: In the formula (m * n * (x) is the closest query point. q pixels, and Indicates in (m * n * Enhanced low-resolution features at positions (m, n), φ k ′ and φ v ′ represents the key network and the value network, respectively. This represents the enhanced low-resolution feature map. Non-local features captured in the process.

8. The method for improving the texture resolution consistency of a 3D model based on arbitrary scale super-resolution as described in claim 5, characterized in that: In step 3.4, the super-resolution image t SR The calculation is shown in the following formula: in, The graph shows the result of bicubic interpolation, f θ (z x x) is an implicit image function, implemented using a multilayer perceptron, z x This represents the implicit encoding corresponding to the x-coordinate of each pixel in the super-resolution image.

9. The method for improving the texture resolution consistency of a 3D model based on arbitrary-scale super-resolution as described in claim 5, characterized in that: In step 3.5, the loss function is: The norm constrains the differences between the super-resolution image and the high-resolution image at the pixel level.

10. A system for improving the texture resolution consistency of 3D models based on arbitrary-scale super-resolution, characterized in that: It includes a processor and a memory, the memory being used to store program instructions, and the processor being used to call the stored instructions in the memory to execute the method for improving the texture resolution consistency of a 3D model based on arbitrary scale super-resolution as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Three-dimensional reconstruction scene texture display method

    CN108335357A

  • Three-dimensional seafloor model establishment method for project sea and calculation method of marine organism loss amount

    CN108665549A