Three-dimensional object geometric attribute optimization method

Through the geometric optimization method based on diffusion model, the geometric distortion and detail loss problems of three-dimensional reconstruction technology in complex scenarios are solved, and high-precision three-dimensional model generation and optimization are achieved, improving the realistic and applicability of the model.

CN120107489APending Publication Date: 2025-06-06BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510277104.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

When existing three-dimensional reconstruction technologies deal with complex, dynamic or detailed scenarios, there are problems such as geometric distortion, surface unsmoothing, details loss or local missing, and traditional geometric optimization methods have limited effects on global optimization.

Method used

The geometric optimization method based on the diffusion model is adopted, and by generating high-precision normal maps and three-dimensional models, multi-view information is used for global optimization, and data blank areas are automatically filled and missing geometric details are restored.

Benefits of technology

It realizes efficient optimization based on low-quality input data, and the generated three-dimensional model is more accurate in geometric structure and details, improving the realistic and applicability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107489A_ABST
    Figure CN120107489A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a three-dimensional object geometric attribute optimization method. A specific embodiment of the method comprises the steps of generating a three-dimensional geometric model data set based on a preset initial three-dimensional model data set and preset rendering angle information; determining three-dimensional geometric model data meeting a preset model data condition in the three-dimensional geometric model data set as a primary three-dimensional geometric model; generating each primary normal map and each primary rendering map based on the primary three-dimensional geometric model; generating feature information of each image based on each primary normal map, each primary rendering map and a preset condition image; performing data decoding processing on each piece of image feature information to obtain each target normal map; and generating a target three-dimensional geometric model based on each target normal map. According to the embodiment, the precision of the generated three-dimensional model can be improved, and the three-dimensional model which is more accurate in geometric structure and details can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the field of computer technology, and in particular to a method for optimizing geometric properties of three-dimensional objects. Background Art

[0002] With the rapid development of computer vision and artificial intelligence technology, 3D reconstruction technology has been widely used in many fields, promoting the development of virtual reality (VR), augmented reality (AR), urban modeling, medical imaging, film and television production, unmanned driving and other industries. Virtual reality and augmented reality technologies require the generation of high-precision, detailed 3D models to provide immersive and interactive user experiences. In terms of urban modeling, 3D reconstruction can provide accurate digital 3D models for urban planning and smart city construction, which is helpful for spatial analysis, simulation and environmental monitoring. In the field of medical imaging, 3D reconstruction technology can help doctors generate 3D models of lesions by processing images such as CT (Computed Tomography) and magnetic resonance imaging, thereby providing more accurate diagnosis and treatment plans. In addition, fields such as film and television production, game design, and autonomous driving also rely on 3D reconstruction technology to generate highly realistic scenes and objects.

[0003] However, despite the significant progress made in these applications, the generated 3D models in reality still face many technical problems, especially in the modeling process of complex scenes and objects with rich details. Current 3D reconstruction methods, especially image-based reconstruction methods, usually rely on depth maps or multi-view images to restore the 3D space. These methods can produce relatively accurate reconstruction results in simple scenes, but there are still obvious quality bottlenecks when dealing with complex, dynamic or detailed scenes. Especially in the reconstruction of complex urban street scenes, indoor environments or natural landscapes, traditional 3D reconstruction technology is easily affected by factors such as noise, perspective problems, lighting changes and object occlusion, resulting in the final generated model having problems such as geometric distortion, rough surface, loss of details or local missing.

[0004] First, the problem of geometric distortion in 3D reconstruction is particularly prominent. Since images from different perspectives may be affected by factors such as lens distortion and lighting changes during the shooting process, the generated mesh model often cannot accurately describe the true geometric shape of the object. In addition, existing reconstruction methods often have shortcomings in detail recovery, especially in the reconstruction of objects with complex surface textures or small scales, where details are easily lost or reconstructed incompletely.

[0005] Secondly, the impact of missing data and noise is also a major challenge in 3D reconstruction. Most 3D reconstruction methods rely on image or depth map data collected by sensors. However, these sensors often cannot cover all the details of the scene when capturing data, especially when objects are partially obscured or the field of view is incomplete, which often results in missing data. At the same time, the noise problem of the sensor is also difficult to avoid. The presence of noise not only affects the accuracy of the data, but also may lead to geometric instability of the reconstruction results, further affecting the quality of the final model.

[0006] In traditional geometry optimization methods, mesh reconstruction or surface smoothing techniques are usually used to repair defects in the data. The surface quality of the model can be partially improved through post-processing steps, such as Laplace smoothing or surface reconstruction algorithms. However, these methods can only locally solve the surface smoothness problem and have limited global optimization effects for large-scale complex scenes. Especially in areas where there is a lack of sufficient data, traditional methods may not be able to recover the lost details, resulting in the integrity of the model being affected.

[0007] In this context, the geometric optimization method based on the diffusion model has gradually become a promising research direction. As an advanced generative model, the diffusion model has been widely used in image processing, image generation, and three-dimensional shape optimization in recent years. Its core idea is to gradually guide the data to approach the target state by simulating a generation method based on the physical diffusion process. In the field of geometric optimization, the diffusion model can automatically fill in the blank areas in the data, restore the missing geometric details, and repair the rough or distorted areas of the surface by learning the correlation between the data. Unlike traditional geometric optimization methods, the diffusion model does not rely on manually designed features or constraints during the optimization process, but performs adaptive repair in a data-driven manner, which can handle complex geometric shapes more accurately.

[0008] In summary, the diffusion model-based geometric optimization method proposed in this paper provides a potential solution to address the shortcomings of traditional 3D reconstruction methods, especially in the reconstruction of complex scenes, dynamic objects and the handling of data missing problems, showing strong advantages.

[0009] Currently, the following method is used to optimize 3D models: First, the input mesh model is converted into a differentiable rendering form (e.g., mesh). Then, the model is optimized using the SDS (Score Distillation Sampling) method. In order to ensure the surface smoothness of the mesh during the optimization process, the Laplace loss is used as a supervisory signal. Finally, the mesh is re-meshed regularly. Finally, the optimized mesh model is converted back to a standard mesh representation (e.g., a triangular mesh) to obtain the result after geometry optimization.

[0010] However, when the above method is used to optimize the 3D model, the following technical problems often occur:

[0011] First, the generated 3D model is too smooth and many details are lost. Although this method can generate a rough 3D structure, the final generated model is often too smooth due to the lack of accurate capture of tiny details. Many details that should exist (such as tiny bumps and texture differences on the surface of the object) will be ignored or simplified. This over-smoothing effect makes the generated 3D model lack realism and layering, especially in scenes or objects that need to be displayed in detail (such as human faces, complex architectural textures, etc.).

[0012] Second, the accuracy of differentiable rendering is limited, and some more sophisticated models may degrade after input. In differentiable rendering technology, due to the limitation of its expressive power, detailed 3D models may not maintain their original accuracy and complexity after being processed by the network. Especially when dealing with highly detailed or delicate objects, the model is prone to degradation, resulting in loss of details or morphological deformation.

[0013] Third, there is a trade-off problem when optimizing texture and geometry at the same time, which affects the overall quality. In the process of optimizing 3D models, the optimization of texture and geometry is usually an interdependent process. If both are optimized at the same time, one optimization direction may have an adverse effect on the other direction. For example, the fineness of the geometric structure may be sacrificed during texture optimization, while the realism of the texture may be ignored during geometry optimization, resulting in a compromise in both texture and geometry in the final model. This trade-off in the optimization process makes it difficult for the model to achieve an ideal state in terms of accuracy. Summary of the invention

[0014] The content of this disclosure is used to introduce concepts in a brief form, which will be described in detail in the detailed implementation section below. The content of this disclosure is not intended to identify the key features or essential features of the technical solution claimed for protection, nor is it intended to limit the scope of the technical solution claimed for protection.

[0015] Some embodiments of the present disclosure propose a method for optimizing geometric properties of three-dimensional objects to solve one or more of the technical problems mentioned in the above background technology section.

[0016] In a first aspect, some embodiments of the present disclosure provide a method for optimizing geometric properties of a three-dimensional object, the method comprising: generating a three-dimensional geometric model data set based on a preset initial three-dimensional model data set and preset rendering angle information; determining the three-dimensional geometric model data in the three-dimensional geometric model data set that meets the preset model data conditions as a primary three-dimensional geometric model; generating primary normal maps and primary renderings based on the primary three-dimensional geometric model, wherein each of the primary normal maps corresponds to camera pose information, and each of the primary renderings corresponds to camera pose information; generating image feature information based on the primary normal maps, the primary renderings and preset conditional images; performing data decoding processing on the image feature information to obtain target normal maps; and generating a target three-dimensional geometric model based on the target normal maps.

[0017] The present invention has been verified on our self-built dataset. By inputting multi-view rough images and normal maps, combined with the target image as a condition, it successfully realizes the generation of high-precision normal maps from low-quality geometric models, and generates high-quality three-dimensional models through the SDF reconstruction algorithm. Thanks to the diffusion model and multi-view normal map optimization technology adopted by the present invention, our method has extremely low requirements for input data, and can effectively reconstruct directly from low-quality images and normal maps, while using external condition images to further improve the accuracy and details of the normal map. Compared with traditional methods, the present invention can not only perform efficient optimization based on low-quality input, but also perform global optimization through multi-view information to ensure that the final generated three-dimensional model is more accurate in geometric structure and details. This method has significant advantages, especially in applications in the fields of virtual reality, augmented reality, game development and three-dimensional reconstruction, which can greatly improve the realism and applicability of the model. In addition, the generated high-quality three-dimensional model is of great significance for data-driven scene understanding, object recognition and spatial analysis research, and at the same time improves the generalization ability of the model in different scenarios. In this way, our technology provides a more efficient and accurate solution for high-quality three-dimensional scene reconstruction and application. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.

[0019] Figure 1 Flow charts of some embodiments of the method for optimizing geometric properties of three-dimensional objects according to the present disclosure;

[0020] Figure 2 It is a schematic diagram of a process for generating a target normal map according to the three-dimensional object geometric property optimization method disclosed in the present invention. DETAILED DESCRIPTION

[0021] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0022] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other.

[0023] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0024] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0025] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0026] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0027] Figure 1 The process 100 of some embodiments of the method for optimizing geometric properties of a three-dimensional object according to the present disclosure is shown. The method for optimizing geometric properties of a three-dimensional object comprises the following steps:

[0028] Step 101: Generate a 3D geometric model dataset based on a preset initial 3D model dataset and preset information of various rendering angles.

[0029] In some embodiments, the execution subject (e.g., a computing device) of the three-dimensional object geometric property optimization method may generate a three-dimensional geometric model data set based on a preset initial three-dimensional model data set and preset rendering angle information. Among them, each initial three-dimensional model data in the above-mentioned initial three-dimensional model data set may be a three-dimensional model corresponding to an object. Each rendering angle information in the above-mentioned rendering angle information may be the Euler angle corresponding to the rendering of the initial three-dimensional model data in the initial three-dimensional model data set. The above-mentioned rendering angle information may include a yaw angle, a pitch angle, and a roll angle. For example, the above-mentioned rendering angle information may be (0 degrees, 0 degrees, 0 degrees). Each three-dimensional geometric model data in the above-mentioned three-dimensional geometric model data set may be the processed initial three-dimensional model data. The above-mentioned execution subject may be a server.

[0030] In some optional implementations of some embodiments, the execution subject may generate a 3D geometric model dataset based on a preset initial 3D model dataset and preset information of various rendering angles through the following steps:

[0031] In the first step, for each initial 3D model data in the initial 3D model data set, the following steps are performed:

[0032] In the first sub-step, for each of the above-mentioned rendering angle information, based on the above-mentioned rendering angle information, the above-mentioned initial three-dimensional model data is subjected to model rendering processing to obtain a model rendering image. The above-mentioned model rendering image may be a rendering image of the above-mentioned initial three-dimensional model data at a corresponding rendering angle. The above-mentioned rendering angle may be an angle represented by the above-mentioned rendering angle information. In practice, the above-mentioned execution entity may input the above-mentioned rendering angle information and the above-mentioned initial three-dimensional model data into a model rendering software to obtain a model rendering image. The above-mentioned model rendering software may be software capable of rendering a three-dimensional model. For example, the above-mentioned model rendering software may be Blender.

[0033] The second sub-step is to generate each model noise rendering image based on each model rendering image obtained. Among them, each model noise rendering image in the above-mentioned model noise rendering images can be a model rendering image after adding noise. In practice, the above-mentioned execution subject can randomly add Gaussian noise of different intensities to the above-mentioned model rendering images through a library function to obtain each model noise rendering image. Among them, the above-mentioned library function can be a function that can add Gaussian noise to an image. For example, the above-mentioned library function can be OpenCV.

[0034] The third sub-step is to perform model reconstruction processing on the above-mentioned model noise renderings to obtain voxel grid models corresponding to the above-mentioned model noise renderings. Among them, the above-mentioned voxel grid model can be a three-dimensional voxel grid corresponding to the above-mentioned model noise renderings. In practice, first, the above-mentioned execution subject can convert the above-mentioned model noise renderings into corresponding 3D Gaussian distributions through a model reconstruction algorithm. Among them, the above-mentioned model reconstruction algorithm can be an algorithm that can generate corresponding 3D Gaussian distributions according to each model rendering. For example, the above-mentioned model reconstruction algorithm can be a GRM algorithm (Generalized Relation Modeling). Then, the above-mentioned execution subject can extract the corresponding signed distance field from the above-mentioned 3D Gaussian distribution through a three-dimensional model construction function. Among them, the above-mentioned three-dimensional model construction function can be a TSDF function (Truncated Signed Distance Function). The above-mentioned signed distance field can be a TSDF field. Finally, the above-mentioned execution subject can extract a voxel grid model from the above-mentioned TSDF field through a grid model extraction algorithm. Among them, the above-mentioned grid model extraction algorithm can be a Marching Cubes algorithm.

[0035] The fourth sub-step is to generate three-dimensional geometric model data corresponding to the voxel grid model based on the voxel grid model. In practice, the execution subject may determine the voxel grid model as three-dimensional geometric model data.

[0036] In the second step, each generated three-dimensional geometric model data is determined as a three-dimensional geometric model data set.

[0037] Step 102: determine the three-dimensional geometric model data satisfying the preset model data condition in the three-dimensional geometric model data set as the primary three-dimensional geometric model.

[0038] In some embodiments, the execution entity may determine the three-dimensional geometric model data that meets the preset model data condition in the three-dimensional geometric model data set as a primary three-dimensional geometric model. The preset model data condition may be that the three-dimensional geometric model data is any one of the three-dimensional geometric model data sets. The primary three-dimensional geometric model includes various vertices and various triangular facets. Each of the triangular facets may be a triangular area composed of three vertices in each of the vertices included in the primary three-dimensional geometric model.

[0039] Step 103, generating each primary normal map and each primary rendering image based on the primary three-dimensional geometric model.

[0040] In some embodiments, the execution subject may generate each primary normal map and each primary rendering based on the primary three-dimensional geometric model. Among them, each primary normal map in the primary normal maps may be a low-quality normal map corresponding to the primary three-dimensional geometric model. Each primary normal map corresponds to different angles of the primary three-dimensional geometric model. Each primary normal map in the primary normal maps corresponds to camera pose information. Each primary rendering in the primary renderings may be a low-quality rendering corresponding to the primary three-dimensional geometric model. Each primary rendering corresponds to different angles of the primary three-dimensional geometric model. Each primary rendering in the primary renderings corresponds to camera pose information. Among them, the camera pose information may be information corresponding to the camera pose. The camera pose information may include an intrinsic parameter matrix, a rotation matrix, a translation vector, a horizontal rotation angle, and a pitch angle. The horizontal rotation angle may represent the rotation angle of the primary three-dimensional geometric model in the horizontal direction. The pitch angle may represent the rotation angle of the primary three-dimensional geometric model in the vertical direction.

[0041] In some optional implementations of some embodiments, the execution subject may generate each primary normal map and each primary rendering image based on the primary three-dimensional geometric model through the following steps:

[0042] The first step is to generate plane normal information corresponding to each of the triangular facets included in the above-mentioned primary three-dimensional geometric model. Each of the above-mentioned triangular facets corresponds to camera pose information. The above-mentioned plane normal information can be a normal vector corresponding to the above-mentioned triangular facet. The above-mentioned plane normal information corresponds to camera pose information. In practice, first, the above-mentioned execution subject can import the above-mentioned primary three-dimensional geometric model into a model building platform. The above-mentioned model building platform can be a platform that can build a model. For example, the above-mentioned model building platform can be Unity. Secondly, the coordinates corresponding to each vertex in the above-mentioned primary three-dimensional geometric model can be obtained as each vertex coordinate through the vertex acquisition component. The above-mentioned vertex acquisition component can be a component that can obtain the vertex coordinates of the model. For example, the above-mentioned vertex acquisition component can be the vertices attribute of the mesh component in Unity. Then, for each of the above-mentioned triangular facets, the coordinates of each vertex corresponding to the above-mentioned triangular facet can be determined as each target vertex coordinate. Finally, the normal vector corresponding to each target vertex coordinate can be determined as the plane normal information corresponding to the above-mentioned triangular facet.

[0043] The second step is to normalize the generated plane normal information to obtain each normalized normal information. Among them, each normalized normal information in the above normalized normal information can be plane normal information that has been normalized. Each normalized normal information in the above normalized normal information corresponds to camera pose information. In practice, for each plane normal information in the above plane normal information, the execution subject can normalize the plane normal information to obtain the normalized plane normal information as the normalized normal information.

[0044] In the third step, for each of the preset camera pose information, perform the following steps:

[0045] In the first sub-step, each normalized normal information corresponding to the camera pose information in the above-mentioned normalized normal information is determined as each target normal information. In practice, each normalized normal information corresponding to the camera pose information that is the same as the camera pose information in the above-mentioned normalized normal information can be determined as each target normal information.

[0046] The second sub-step is to perform coordinate transformation processing on the above-mentioned target normal information based on the above-mentioned camera posture information to obtain each target normal coordinate information. Among them, each target normal coordinate information in the above-mentioned target normal coordinate information can be the target normal information after coordinate transformation. Each target normal coordinate information in the above-mentioned target normal coordinate information can include a horizontal coordinate value, a vertical coordinate value and a vertical coordinate value. In practice, for each target normal information in the above-mentioned target normal information, the above-mentioned execution entity can determine the product of the rotation matrix included in the above-mentioned camera posture information and the above-mentioned target normal information as the target normal coordinate information.

[0047] In the third sub-step, for each target normal coordinate information in the above-mentioned target normal coordinate information, data conversion processing is performed on the above-mentioned target normal coordinate information to obtain color channel information corresponding to the above-mentioned target normal coordinate information. Among them, the above-mentioned color channel information can be the RGB value corresponding to the target normal coordinate information. In practice, for each target normal coordinate information in the above-mentioned target normal coordinate information, first, the above-mentioned execution subject can determine the sum of the horizontal coordinate value included in the above-mentioned target normal coordinate information and the first preset coordinate value as the normal horizontal coordinate value. Among them, the above-mentioned first preset coordinate value can be a preset value. Here, the specific setting of the above-mentioned first preset coordinate value is not limited. For example, the above-mentioned first preset coordinate value can be 1. Then, the ratio of the above-mentioned normal horizontal coordinate value to the second preset coordinate value can be determined as the red color channel value. Among them, the above-mentioned second preset coordinate value can be a preset value. Here, the specific setting of the above-mentioned second preset coordinate value is not limited. For example, the above-mentioned second preset coordinate value can be 2. Then, the execution subject may generate a green color channel value and a blue color channel value based on the ordinate value and the vertical coordinate value included in the target normal coordinate information. The green color channel value and the blue color channel value may refer to the specific implementation method of generating the red color channel value, which will not be described in detail here. Finally, the red color channel value, the green color channel value and the blue color channel value may be combined into color channel information.

[0048] The fourth sub-step is to generate a primary normal map corresponding to the camera pose information based on the preset two-dimensional texture image and the obtained color channel information. The two-dimensional texture image may be a blank two-dimensional image corresponding to the primary three-dimensional geometric model. For example, the two-dimensional texture image may be an image whose pixel values ​​are all 255. The two-dimensional texture image corresponds to width data and height data. In practice, first, each vertex corresponding to the camera pose information included in the primary three-dimensional geometric model is determined as a target vertex. Then, the execution subject may obtain the two-dimensional texture coordinates corresponding to each target vertex in the target vertices through a two-dimensional texture coordinate acquisition function. The two-dimensional texture coordinates may be UV coordinates corresponding to the two-dimensional texture image. The two-dimensional texture coordinates may include horizontal coordinate values ​​and vertical coordinate values. The horizontal coordinate value may be a coordinate value in the horizontal direction of the two-dimensional texture image. The vertical coordinate value may be a coordinate value in the vertical direction of the two-dimensional texture image. The two-dimensional texture coordinate acquisition function may be a function that can obtain the two-dimensional texture coordinates corresponding to the model vertex. For example, the two-dimensional texture coordinate acquisition function may be the Mesh.GetUVs function in Unity. Then, for each color channel information in the above-mentioned color channel information, the triangular face corresponding to the above-mentioned color channel information can be determined as the target triangular face. Then, each two-dimensional texture coordinate corresponding to the above-mentioned target triangular face in each of the obtained two-dimensional texture coordinates can be determined as each target two-dimensional texture coordinate. Then, for each of the above-mentioned target two-dimensional texture coordinates, the product of the horizontal coordinate value included in the above-mentioned target two-dimensional texture coordinate and the width data corresponding to the above-mentioned two-dimensional texture image can be determined as the image horizontal coordinate value. Then, the above-mentioned image horizontal coordinate value can be input into a floor function to obtain a pixel horizontal coordinate value. Among them, the above-mentioned floor function can be a function that rounds down the data. For example, the above-mentioned floor function can be a floor function. Then, the product of the vertical coordinate value included in the above-mentioned target two-dimensional texture coordinate and the height data corresponding to the above-mentioned two-dimensional texture image can be determined as the image vertical coordinate value. Then, the above-mentioned image vertical coordinate value can be input into the above-mentioned floor function to obtain a pixel vertical coordinate value. Then, the above-mentioned pixel horizontal coordinate value and the above-mentioned pixel vertical coordinate value can be combined into a pixel coordinate. Then, the pixel point corresponding to the pixel coordinate in the two-dimensional texture image can be determined as the target pixel point. Then, the color channel information can be determined as the RGB value corresponding to the pixel point to fill the two-dimensional texture image. Finally, the two-dimensional texture image with all the target vertices filled is determined as the primary normal map.

[0049] The fifth sub-step is to generate color data corresponding to the target normal information for each of the target normal information, based on the preset illumination information and the target normal information. The illumination information may be information corresponding to the illumination simulation of the primary three-dimensional geometric model. The illumination information may include light source position information, ambient light intensity information and diffuse reflection coefficient information. The light source position information may be a vector for characterizing the illumination direction of the light source. The ambient light intensity information may characterize the intensity of light in the virtual space where the primary three-dimensional geometric model is located. The diffuse reflection coefficient information may characterize the reflection intensity when the primary three-dimensional geometric model diffusely reflects light. The color data may be RGB values ​​corresponding to the target normal information.

[0050] The sixth sub-step is to generate a primary rendering corresponding to the above-mentioned camera pose information based on the generated color data.

[0051] In some optional implementations of some embodiments, the execution subject may generate color data corresponding to the target normal information based on preset illumination information and the target normal information through the following steps:

[0052] In the first step, the dot product between the light source position information and the target normal information is determined as reflection data.

[0053] In the second step, in response to determining that the reflection data satisfies a preset reflection data condition, the reflection data is determined as target reflection data. The preset reflection data condition may be that the reflection data is greater than a preset standard reflection data. The standard reflection data may be a preset value. Here, the specific setting of the standard reflection data is not limited.

[0054] In a third step, in response to determining that the reflection data does not satisfy the preset reflection data condition, the preset standard reflection data is determined as the target reflection data.

[0055] The fourth step is to determine the product of the diffuse reflection coefficient information and the target reflection data as diffuse reflection data.

[0056] In the fifth step, the sum of the ambient light intensity information and the diffuse reflection data is determined as the illumination color data.

[0057] Step 6: Determine color data based on the illumination color data. In practice, the execution subject may combine three values ​​that are the same as the values ​​corresponding to the illumination color data into color data. For example, when the illumination color data is 1, the corresponding color data is (1, 1, 1).

[0058] Step 104, generating each image feature information based on each primary normal map, each primary rendering image and a preset conditional image.

[0059] In some embodiments, the execution subject may generate each image feature information based on each primary normal map, each primary rendering image and a preset conditional image, wherein the conditional image may be an image of a real object corresponding to the primary three-dimensional geometric model.

[0060] In the process of adopting technical solutions to solve the above technical problems, the following problems are often accompanied:

[0061] The methods such as Laplace loss and re-gridding operations used in the optimization process tend to simplify the geometric details of the surface of the three-dimensional model, which in turn easily leads to lower precision of the optimized three-dimensional model, resulting in the need to consume computing resources to re-optimize the three-dimensional model, resulting in a waste of computing resources.

[0062] Faced with the above technical problems, we decided to adopt the following solutions:

[0063] In some optional implementations of some embodiments, the execution subject may generate each image feature information based on each primary normal map, each primary rendering image and a preset conditional image through the following steps:

[0064] For each of the primary normal maps above, perform the following steps:

[0065] In the first step, the primary rendering corresponding to the primary normal map in each of the primary renderings is determined as the target primary rendering. In practice, the execution subject may determine the primary rendering whose camera pose information corresponding to each of the primary renderings is the same as the camera pose information corresponding to the primary normal map as the target primary rendering.

[0066] In the second step, feature extraction processing is performed on the primary normal map and the target primary rendering image to obtain primary normal feature information and primary rendering feature information. The primary normal feature information may be a feature vector corresponding to the primary normal map. The primary rendering feature information may be a feature vector corresponding to the target primary rendering image. In practice, the execution entity may input the primary normal map and the target primary rendering image into the encoding layer of a pre-trained diffusion model to obtain primary normal feature information and primary rendering feature information. The diffusion model may be a neural network model that takes the primary normal map, the target primary rendering image and the conditional image as input and takes the image feature information as output. The diffusion model may include eight layers.

[0067] The first layer may be an encoding layer. The encoding layer may be an encoder that takes the primary normal map and the target primary rendering as input and outputs the primary normal feature information corresponding to the primary normal map and the primary rendering feature information corresponding to the target primary rendering. For example, the encoder may be a variational autoencoder.

[0068] The second layer may be a feature concatenation layer. The feature concatenation layer may be a concatenation function that takes the primary normal feature information and the primary rendering feature information as input and takes the feature embedding information as output. For example, the concatenation function may be a concat function. The feature embedding information may be a feature vector obtained by concatenating the primary normal feature information and the primary rendering feature information.

[0069] The third layer may be a feature extraction network. The feature extraction network may be a neural network that takes the conditional image as input and takes the conditional feature information corresponding to the conditional image as output. For example, the feature extraction network may be a convolutional neural network. The conditional feature information may be a feature vector corresponding to the conditional image.

[0070] The fourth layer may be a feature information splicing layer. The feature information splicing layer may be a feature splicing function that takes the feature embedding information output by the feature splicing layer and the conditional feature information output by the feature extraction network as input and takes the input feature information as output. For example, the feature splicing function may be an add function. The input feature information may be a feature vector obtained by splicing the feature embedding information and the conditional feature information.

[0071] The fifth layer may be a feature fusion network. The feature fusion network may be a neural network that takes input feature information as input and takes fused feature information as output. The feature fusion network may be a Transformer module obtained by fine-tuning Stable Video Diffusion. The fused feature information may be a feature vector obtained by optimizing the input feature information. In practice, first, the execution entity may use the labeled image feature vector data set and the mean square error loss function to fine-tune the Stable Video Diffusion. Then, the Transformer module included in the fine-tuned Stable Video Diffusion may be used as a feature fusion network. The image feature vector data in the image feature vector data set may be a feature vector obtained by fusing the feature vectors of each image. The fused feature information may be a feature vector obtained by optimizing the input feature information.

[0072] The sixth layer may be a feature enhancement network. The feature enhancement network may be a neural network that takes fused feature information as input and takes enhanced feature information as output. For example, the feature enhancement network may be a pre-trained feedforward neural network. The pre-training may be a process of fine-tuning the feedforward neural network using a labeled fused feature information set and a mean square error loss function. The enhanced feature information may be a feature vector obtained by processing the fused feature information through a pre-trained feedforward neural network.

[0073] The seventh layer may be a processing layer. The processing layer may include a splicing layer and a normalization layer. The splicing layer may be a splicing function that takes the fused feature information output by the feature fusion network and the enhanced feature information output by the feature enhancement network as input and takes the splicing feature information as output. The splicing function may be an add function. The splicing feature information may be a feature vector obtained by splicing the fused feature information and the enhanced feature information. The normalization layer may be a normalization layer that takes the splicing feature information as input and takes the normalized feature information as output. The normalized feature information may be the splicing feature information processed by the normalization layer.

[0074] The eighth layer may be a feature denoising network. The feature denoising network may be a neural network that takes normalized feature information as input and image feature information as output. For example, the feature denoising network may be a convolutional neural network. The image feature information may be normalized feature information after denoising.

[0075] The third step is to perform feature concatenation processing on the primary normal feature information and the primary rendering feature information to obtain feature embedding information. In practice, the execution subject may input the primary normal feature information and the primary rendering feature information into the feature concatenation layer in the diffusion model to obtain feature embedding information.

[0076] The fourth step is to perform feature extraction processing on the conditional image to obtain conditional feature information corresponding to the conditional image. In practice, the execution subject may input the conditional image into the feature extraction network in the diffusion model to obtain the conditional feature information.

[0077] Optionally, the execution subject may also adjust the conditional image by an image processing algorithm so that the size of the adjusted conditional image is the same as that of the primary normal map. Figure 1 Then, the adjusted condition image can be input into the feature extraction network in the diffusion model to obtain condition feature information.

[0078] Step 5: perform feature concatenation processing on the feature embedding information and the conditional feature information to obtain input feature information. In practice, the execution entity may input the feature embedding information and the conditional feature information into the feature information concatenation layer in the diffusion model to obtain input feature information.

[0079] Step 6: Based on the input feature information, perform the following iterative steps:

[0080] The first sub-step is to update the number of iterations based on a preset value. The preset value may be a pre-set value. The specific setting of the preset value is not limited. The number of iterations may be a value used to characterize the number of executions of the iterative step. In practice, the execution subject may determine the sum of the number of iterations and the preset value as the number of iterations to update the number of iterations.

[0081] The second sub-step is to generate fused feature information based on the input feature information. In practice, the execution subject may input the input feature information into the feature fusion network in the diffusion model to obtain fused feature information.

[0082] The third sub-step is to perform feature enhancement processing on the fused feature information to obtain enhanced feature information. In practice, the execution subject may input the fused feature information into the feature enhancement network in the diffusion model to obtain enhanced feature information.

[0083] The fourth sub-step is to generate normalized feature information based on the fused feature information and the enhanced feature information. In practice, the execution entity may input the fused feature information and the enhanced feature information into a processing layer in the diffusion model to obtain normalized feature information.

[0084] In a fifth sub-step, in response to determining that the updated number of iterations does not meet the preset number of iterations condition, the normalized feature information is determined as the input feature information, and the iterative step is performed again using the updated input feature information. The preset number of iterations condition may be that the updated number of iterations is greater than or equal to a preset target number of iterations. The target number of iterations may be a pre-set value. Here, the specific setting of the target number of iterations is not limited.

[0085] Step 7: Generate image feature information based on the updated normalized feature information. In practice, the execution subject may input the updated normalized feature information into the feature denoising network in the diffusion model to obtain the denoised normalized feature information as the image feature information.

[0086] The above technical solution and its related contents are combined with steps 101 to 106 as an inventive point of an embodiment of the present disclosure, which solves the problem of "waste of computing resources". The factors that lead to waste of computing resources are often as follows: the methods such as Laplace loss and re-gridding operation adopted in the optimization process are easy to simplify the geometric details of the surface of the three-dimensional model, which in turn easily leads to the low fineness of the optimized three-dimensional model, resulting in the need to consume computing resources to re-optimize the three-dimensional model, resulting in waste of computing resources. If the above factors are solved, the waste of computing resources can be reduced. In order to achieve this effect, the present disclosure first performs the following steps for each primary normal map in the above-mentioned primary normal maps: Secondly, the primary rendering corresponding to the above-mentioned primary normal map in the above-mentioned primary renderings is determined as the target primary rendering. Thus, the primary rendering corresponding to the above-mentioned primary normal map can be determined. Then, the above-mentioned primary normal map and the above-mentioned target primary rendering are subjected to feature extraction processing to obtain primary normal feature information and primary rendering feature information. Thus, the feature vector corresponding to the above-mentioned primary normal map and the feature vector corresponding to the above-mentioned target primary rendering can be obtained. Then, feature splicing processing is performed on the primary normal feature information and the primary rendering feature information to obtain feature embedding information. Thus, the primary normal feature information and the primary rendering feature information can be spliced ​​to obtain feature embedding information. Then, feature extraction processing is performed on the conditional image to obtain conditional feature information corresponding to the conditional image. Thus, a feature vector corresponding to the conditional image can be obtained. Then, feature splicing processing is performed on the feature embedding information and the conditional feature information to obtain input feature information. Thus, input feature information can be obtained. Then, based on the input feature information, the following iterative steps are performed: first, based on a preset value, the number of iterations is updated. Thus, the number of executions of the iterative steps can be recorded. Then, based on the input feature information, fused feature information is generated. Thus, the input feature information can be further processed to obtain fused feature information. Then, feature enhancement processing is performed on the fused feature information to obtain enhanced feature information. Thus, enhanced feature information can be obtained. Then, based on the fused feature information and the enhanced feature information, normalized feature information is generated. Thus, normalized feature information can be obtained. Then, in response to determining that the updated number of iterations does not meet the preset number of iterations condition, the normalized feature information is determined as the input feature information, and the iterative step is performed again using the updated input feature information. Thus, the iterative step can be performed again using the updated input feature information. Finally, based on the updated normalized feature information, the image feature information is generated. Thus, the image feature information can be obtained.Also, when optimizing the primary three-dimensional geometric model, a conditional image can be introduced to provide texture details needed for reference in the optimization process, and then the conditional image, the primary normal maps and the primary renderings are processed to obtain image feature information integrated with the conditional image. Finally, the primary three-dimensional geometric model is optimized through the image feature information, thereby improving the fineness of the optimized primary three-dimensional geometric model, thereby reducing the probability of needing to re-optimize due to the low fineness of the three-dimensional model, thereby reducing the waste of computing resources.

[0087] Step 105, performing data decoding processing on each image feature information to obtain each target normal map.

[0088] In some embodiments, the execution subject may perform data decoding processing on the above-mentioned image feature information to obtain each target normal map. Among them, each of the above-mentioned target normal maps may be a high-quality normal map corresponding to the primary normal map. In practice, the execution subject may input the above-mentioned image feature information into a decoder to obtain each target normal map. Among them, the above-mentioned decoder may decode the feature vector into a normal map. For example, the above-mentioned decoder may be a decoder in a variational autoencoder.

[0089] Optionally, the execution subject may further perform denoising on the conditional image by an image denoising algorithm to obtain the denoised conditional image as the denoised conditional image. For example, the image denoising algorithm may be a median filter algorithm. Then, the image feature information and the denoised conditional image may be input into the decoder, so that the decoder can generate corresponding target normal maps with reference to the denoised conditional image.

[0090] Figure 2 FIG. 1 is a flow chart of generating a target normal map according to the three-dimensional object geometric property optimization method disclosed in the present invention. Figure 2 As shown, Figure 2 The rough image in may be the target primary rendering. The rough normal map may be a primary normal map corresponding to the target primary rendering. The “first feature splicing symbol” may represent the process of splicing the primary normal feature information corresponding to the primary normal map and the primary rendering feature information corresponding to the target primary rendering to obtain feature embedding information. T Information can be embedded into the features obtained after concatenation. Figure 2The conditional image in can be the above conditional image. The “first image processing symbol” can represent the process of adjusting the above conditional image to obtain the adjusted conditional image. The “second feature splicing” can represent the process of splicing the above feature embedding information and the conditional feature information corresponding to the above conditional image to obtain the input feature information. Figure 2 The "Denoising U-Net" in can represent the process of processing the above input feature information to obtain image feature information. The "Denoising U-Net" can be the above diffusion model. The Q in the "Denoising U-Net" can be the query vector corresponding to the feature fusion network of the fifth layer in the "Denoising U-Net". The K in the "Denoising U-Net" can be the key vector corresponding to the feature fusion network of the fifth layer in the "Denoising U-Net". The V in the "Denoising U-Net" can be the value vector corresponding to the feature fusion network of the fifth layer in the "Denoising U-Net". Z T-1 A can be the obtained image feature information. T-1 It can be an intermediate feature generated in the process corresponding to the above-mentioned "denoising U-Net". The "second image processing symbol" can be a process of denoising the conditional image to obtain a denoised conditional image. The super-resolution optimized normal map can be a target normal map corresponding to the above-mentioned target primary rendering image and the primary normal map.

[0091] Step 106: Generate a target three-dimensional geometric model based on each target normal map.

[0092] In some embodiments, the execution entity may generate a target three-dimensional geometric model based on the target normal maps.

[0093] In the process of adopting technical solutions to solve the above technical problems, the following problems are often accompanied:

[0094] The method of directly converting the three-dimensional model into a differentiable rendering form and then optimizing the converted three-dimensional model involves relatively complex processing steps, consumes a large amount of computing resources during processing, and has low data processing efficiency when a large number of three-dimensional models need to be optimized.

[0095] Faced with the above technical problems, we decided to adopt the following solutions:

[0096] In some optional implementations of some embodiments, the execution subject may generate a target three-dimensional geometric model based on the target normal maps through the following steps:

[0097] The first step is to generate distance field data corresponding to each target normal map based on each target normal map, wherein the distance field data may be a SDF (Signed Distance Field) field corresponding to each pixel normal information.

[0098] In practice, for each target normal map in each of the target normal maps, for each pixel in each of the pixels included in the target normal map, first, the execution subject may determine the pixel as a pixel to be processed. Then, the target normal map corresponding to the pixel to be processed may be determined as the target normal map to be processed. Then, the coordinate position of the pixel to be processed in the target normal map to be processed may be determined as the coordinate of the pixel to be processed. Among them, the coordinate position of the pixel to be processed may be the coordinate determined in the image coordinate system corresponding to the target normal map to be processed. The image coordinate system may be constructed with the first point in the lower left corner of the target normal map to be processed as the pixel origin, the vertical upward direction as the ordinate axis, and the horizontal right direction as the abscissa axis. Then, the preset pixel coordinate value may be added to the pixel coordinate to be processed to obtain the target pixel vector. Among them, the target pixel vector may be the vector corresponding to the pixel coordinate to be processed. The preset pixel coordinate value may be a preset value. For example, the preset pixel coordinate value may be 1. As an example: when the coordinates of the above-mentioned pixel point to be processed are (3,2), the target pixel point vector to which the preset pixel point coordinate values ​​are added is (3,2,1). Then, the transpose of the above-mentioned target pixel point vector can be determined as the target pixel point transpose vector. Then, the camera pose information corresponding to the above-mentioned target normal map to be processed can be determined as the target camera pose information. Then, the intrinsic parameter matrix in the above-mentioned target camera pose information can be determined as the target intrinsic parameter matrix. Then, the product of the inverse matrix of the above-mentioned target intrinsic parameter matrix and the above-mentioned target pixel point transpose vector can be determined as the pixel point direction vector. Then, the product of the above-mentioned pixel point direction vector and the preset depth value can be determined as the pixel point depth vector. Among them, the above-mentioned preset depth value can be a pre-set value. Here, the specific setting of the above-mentioned preset depth value is not limited. Then, the translation vector in the above-mentioned target camera pose information can be determined as the target translation vector. The difference between the above-mentioned pixel point depth vector and the above-mentioned target translation vector can be determined as the pixel point translation vector. Then, the inverse of the above-mentioned pixel normal information can be determined as the pixel inverse normal information. Then, the product of the above pixel inverse normal information and the above pixel translation vector can be determined as a three-dimensional point coordinate. Then, the determined three-dimensional point coordinates can be combined into an array as a three-dimensional point coordinate array. Finally, the above execution subject can obtain distance field data by inputting the above three-dimensional point coordinate array into a three-dimensional model reconstruction function. Among them, the above three-dimensional model reconstruction function can be a function that can model a three-dimensional object. For example, the above three-dimensional model reconstruction function can be a signed distance function.

[0099] The second step is to smooth the distance field data to obtain smoothed distance field data. The smoothed distance field data may be distance field data after smoothing. In practice, the execution subject may smooth the distance field data by an interpolation algorithm to obtain smoothed distance field data. The interpolation algorithm may be trilinear interpolation.

[0100] The third step is to perform data extraction processing on the above-mentioned smooth distance field data to obtain a smooth three-dimensional geometric model corresponding to the above-mentioned smooth distance field data. The above-mentioned smooth three-dimensional geometric model can be a three-dimensional mesh model extracted from the above-mentioned smooth distance field data. In practice, the above-mentioned execution subject can extract the smooth three-dimensional geometric model from the above-mentioned smooth distance field data through a model extraction algorithm. The above-mentioned model extraction algorithm can be a Marching Cubes algorithm.

[0101] The fourth step is to generate a target three-dimensional geometric model corresponding to the smooth three-dimensional geometric model based on the smooth three-dimensional geometric model. The target three-dimensional geometric model may be an optimized smooth three-dimensional geometric model. In practice, the execution subject may perform denoising on the smooth three-dimensional geometric model by a mesh smoothing algorithm to obtain a target three-dimensional geometric model. The mesh smoothing algorithm may be a Laplacian algorithm.

[0102] The above technical solution and its related contents, combined with steps 101 to 106, are an inventive point of an embodiment of the present disclosure, which solves the problem of "large consumption of computing resources". The factors that lead to large consumption of computing resources and low data processing efficiency are often as follows: the method of directly converting a three-dimensional model into a micro-renderable form and then optimizing the converted three-dimensional model involves relatively complex processing steps, consumes large computing resources during processing, and when a large number of three-dimensional models need to be optimized, the efficiency of data processing is low. If the above factors are solved, the consumption of computing resources can be reduced and the efficiency of data processing can be improved. In order to achieve this effect, the present disclosure first generates distance field data corresponding to the above-mentioned each target normal map based on the above-mentioned each target normal map. Thus, the distance field data can be obtained. Secondly, the above-mentioned distance field data is smoothed to obtain smoothed distance field data. Thus, the smoothed distance field data can be obtained. Then, the above-mentioned smoothed distance field data is subjected to data extraction processing to obtain a smoothed three-dimensional geometric model corresponding to the above-mentioned smoothed distance field data. Thus, the smoothed three-dimensional geometric model can be extracted from the above-mentioned smoothed distance field data. Finally, based on the above-mentioned smoothed three-dimensional geometric model, a target three-dimensional geometric model corresponding to the above-mentioned smoothed three-dimensional geometric model is generated. Thus, an optimized primary three-dimensional geometric model can be obtained. Also, each target normal map can be generated by processing each primary normal map, each primary rendering image and conditional image corresponding to the primary three-dimensional geometric model, and finally, the primary three-dimensional geometric model is optimized by each target normal map, so the optimization processing of the three-dimensional model can be converted into the processing of the two-dimensional image, thereby reducing the complex steps involved in the processing, and further reducing the computing resources consumed in the processing, and further improving the processing efficiency.

[0103] The above descriptions are only some preferred embodiments of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) and the technical solutions formed.

Claims

1. A method for optimizing geometric properties of a three-dimensional object, comprising: Generate a 3D geometric model dataset based on a preset initial 3D model dataset and preset information of various rendering angles; Determining the three-dimensional geometric model data that meets the preset model data condition in the three-dimensional geometric model data set as a primary three-dimensional geometric model; Based on the primary three-dimensional geometric model, generating each primary normal map and each primary rendering, wherein each of the primary normal maps corresponds to camera pose information, and each of the primary renderings corresponds to camera pose information; Generate each image feature information based on each primary normal map, each primary rendering image and a preset conditional image; Performing data decoding processing on each image feature information to obtain each target normal map; A target three-dimensional geometric model is generated based on the target normal maps.

2. The method according to claim 1, wherein: The step of generating a 3D geometric model dataset based on a preset initial 3D model dataset and preset rendering angle information includes: For each initial 3D model data in the initial 3D model data set, the following steps are performed: For each of the rendering angle information, based on the rendering angle information, performing model rendering processing on the initial three-dimensional model data to obtain a model rendering image; Based on the obtained rendering images of each model, generating noise rendering images of each model; Performing model reconstruction processing on each of the model noise renderings to obtain a voxel grid model corresponding to each of the model noise renderings; Based on the voxel grid model, generating three-dimensional geometric model data corresponding to the voxel grid model; The generated three-dimensional geometric model data are determined as three-dimensional geometric model data sets.

3. The method according to claim 1, wherein: The primary three-dimensional geometric model includes each vertex and each triangular face; And generating each primary normal map and each primary rendering image based on the primary three-dimensional geometric model includes: For each of the triangular facets included in the primary three-dimensional geometric model, generating plane normal information corresponding to the triangular facet; Normalizing each generated plane normal information to obtain each normalized normal information; For each of the preset camera pose information, perform the following steps: Determining each normalized normal line information corresponding to the camera pose information in the each normalized normal line information as each target normal line information; Based on the camera posture information, coordinate conversion processing is performed on the normal line information of each target to obtain the normal line coordinate information of each target; For each target normal coordinate information in the target normal coordinate information, performing data conversion processing on the target normal coordinate information to obtain color channel information corresponding to the target normal coordinate information; Based on the preset two-dimensional texture image and the obtained information of each color channel, a primary normal map corresponding to the camera pose information is generated; For each target normal information among the respective target normal information, generating color data corresponding to the target normal information based on preset illumination information and the target normal information; Based on the generated color data, a primary rendering corresponding to the camera pose information is generated.

4. The method according to claim 3, wherein: The illumination information includes light source position information, ambient light intensity information and diffuse reflectance information; And the step of generating color data corresponding to the target normal information based on the preset illumination information and the target normal information comprises: Determine the dot product between the light source position information and the target normal information as reflection data; In response to determining that the reflection data satisfies a preset reflection data condition, determining the reflection data as target reflection data; In response to determining that the reflection data does not satisfy the preset reflection data condition, determining preset standard reflection data as target reflection data; Determine the product of the diffuse reflection coefficient information and the target reflection data as diffuse reflection data; Determine the sum of the ambient light intensity information and the diffuse reflection data as illumination color data; Based on the illumination color data, color data is determined.