Surface normal estimation method, device and equipment of target object and medium

By combining flash and non-flash images and using diffusion prior denoising method, inputting a preset network for surface normal estimation, the problem in the prior art is solved that it is difficult to accurately estimate the surface normal of the object under complex lighting conditions, and a higher quality normal estimation and three-dimensional reconstruction effect are achieved.

CN120147395AActive Publication Date: 2025-06-13BEIJING UNIV OF POSTS & TELECOMM +1
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510094535.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-06-13
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

Existing surface normal estimation methods based on deep learning are difficult to deal with objects with rich details and unclear details. Especially under complex lighting conditions, the shape-reflectivity ambiguity is high, making it difficult to accurately estimate the surface normal of the object.

Method used

By combining flash and non-flash images and combining diffusion prior denoising methods, the trained preset network is input for normal estimation, including encoder module, diffusion denoising module and decoder module, reducing shape-reflectivity ambiguity and improving the estimation quality of surface normals.

Benefits of technology

Under complex lighting conditions, it is possible to accurately estimate the surface normal of the target object, especially in the estimation of surface normals with rich details, which improves the quality of the normal map and the accuracy of three-dimensional reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147395A_ABST
    Figure CN120147395A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision, in particular to a surface normal estimation method and device of a target object, equipment and a medium. The method comprises the following steps: acquiring a first image of a target object under a flash lamp condition and a second image of the target object under a non-flash lamp condition; the first image and the second image are preprocessed; inputting the preprocessed image into a trained preset network for normal estimation to obtain a normal graph of the target object; wherein the preset network comprises an encoder module, a diffusion denoising module and a decoder module. According to the method, the surface normal of the target object can be accurately estimated under the complex illumination condition by combining the flash lamp image and the non-flash lamp image and utilizing the diffusion prior denoising method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision technology, and in particular, to a method, apparatus, device, and medium for estimating the surface normal of an object. Background Art

[0002] A surface normal map is a 2.5D geometric representation, and various computer vision tasks, such as 3D shape reconstruction, augmented reality applications, and advanced rendering techniques, require high-quality surface normals. However, creating a surface normal map with rich details usually requires 3D scanning using a high-precision scanner or collecting images with controlled illumination in a darkroom for reconstruction using photometric stereo methods, both of which are cumbersome and require a lot of resources.

[0003] Recently, there have been many advances in surface normal estimation based on deep learning. Most of these methods take a single RGB image as input and estimate the corresponding surface normal map. Although these methods are popular, due to the inherent shape-reflectance ambiguity of single-image input, they usually have difficulty dealing with objects with rich details and ambiguity. Specifically, a single image may be affected by additional shadows that obscure the details of the object, and the ambient lighting conditions during the capture process may not fully reveal the true geometric features of the scene. Summary of the Invention

[0004] To solve the above technical problems, the present application provides a method, apparatus, device, and medium for estimating the surface normal of an object, which improves the estimation quality of the surface normal and reduces the shape-reflectance ambiguity by using flash and non-flash images as well as a diffusion prior, thereby enabling more accurate reconstruction of the surface of an object with rich details.

[0005] In a first aspect of the present application, a method for estimating the surface normal of an object is provided. The method includes: obtaining a first image of the object under flash conditions and a second image of the object under non-flash conditions;

[0006] preprocessing the first image and the second image;

[0007] inputting the preprocessed images into a trained preset network for normal estimation to obtain a normal map of the object;

[0008] wherein the preset network includes an encoder module, a diffusion denoising module, and a decoder module.

[0009] In some embodiments of the present application, the preprocessing of the first image and the second image includes:

[0010] obtaining mask images of the first image and the second image;

[0011] Crop the target object based on the masked image to obtain an image corresponding to the main body part of the target object;

[0012] Calculate and generate a bounding box based on the image corresponding to the main body part to define the area of the main body part, and use the image defined by the bounding box as the preprocessed first image and the preprocessed second image.

[0013] In some embodiments of the present application, the step of inputting the preprocessed image into a trained preset network for normal estimation to obtain the normal map of the target object includes:

[0014] Input the preprocessed first image and the preprocessed second image into the encoder module respectively to extract latent features, and obtain a first feature map corresponding to the first image and a second feature map corresponding to the second image;

[0015] Merge the first feature map, the second feature map and the Gaussian noise map in the channel dimension to obtain a merged feature map;

[0016] Input the merged feature map into the diffusion denoising module and the decoder module in sequence to obtain the normal map of the target object.

[0017] In some embodiments of the present application, the step of inputting the merged feature map into the diffusion denoising module and the decoder module in sequence to obtain the normal map of the target object includes:

[0018] Input the merged feature map into the diffusion denoising module, and the diffusion denoising module gradually removes the noise in the merged feature map to obtain a normal map feature with object surface details;

[0019] Input the normal map feature into the decoder module, and the decoder module restores the normal map feature to the normal map of the target object through deconvolution operation.

[0020] In some embodiments of the present application, the training method of the preset network includes:

[0021] Obtain a synthetic dataset and a real dataset, where the synthetic dataset includes computer-generated images with virtual target objects, and the real dataset includes images of real objects taken;

[0022] Preprocess the synthetic dataset and the real dataset to obtain flash images and non-flash images suitable for input;

[0023] Train the preset network based on the flash images and the non-flash images.

[0024] In some embodiments of the present application, the training method of the preset network further includes:

[0025] When training the diffusion denoising module, Gaussian noise is added to the real normal map to simulate noise interference of different degrees, so that the diffusion denoising module learns to recover the real normal map from the noise.

[0026] In some embodiments of the present application, after obtaining the normal map of the target object, it further includes: performing three-dimensional reconstruction based on the normal map of the target object.

[0027] A second aspect of the present application provides a surface normal estimation device for a target object, the device includes:

[0028] An acquisition module, configured to acquire a first image of the target object under flash conditions and a second image under non-flash conditions;

[0029] A processing module, configured to preprocess the first image and the second image;

[0030] An estimation module, configured to input the preprocessed images into a trained preset network for normal estimation to obtain the normal map of the target object;

[0031] Wherein, the preset network includes an encoder module, a diffusion denoising module, and a decoder module.

[0032] A third aspect of the present application provides an electronic device, including a memory and a processor. When the computer-readable instructions stored in the memory are executed by the processor, the processor executes the surface normal estimation method for the target object in each embodiment of the present application.

[0033] A fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the surface normal estimation method in each embodiment of the present application is implemented.

[0034] The technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:

[0035] The surface normal estimation method of the target object in each embodiment of the present application obtains a first image of the target object under flash conditions and a second image under non-flash conditions, preprocesses the first image and the second image, and inputs the preprocessed images into a trained preset network for normal estimation to obtain the normal map of the target object. Among them, the preset network includes an encoder module, a diffusion denoising module, and a decoder module. Thus, by combining flash and non-flash images and using a denoising method with diffusion prior, this method can accurately estimate the surface normal of the target object under complex lighting conditions, especially having significant advantages in the normal estimation of detailed surfaces.

[0036] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. Brief Description of the Drawings

[0037] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to limit the present application. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0038] Figure 1 is a schematic diagram of the steps of a surface normal estimation method of a target object in an exemplary embodiment of the present application;

[0039] Figure 2 is a schematic diagram of the structure of a preset network in an exemplary embodiment of the present application;

[0040] Figure 3 is a schematic diagram of a photographing device in an exemplary embodiment of the present application;

[0041] Figure 4 is a schematic diagram of a real dataset of photographs in an exemplary embodiment of the present application;

[0042] Figure 5 is a schematic diagram of the process of normal estimation through a preset network in an exemplary embodiment of the present application;

[0043] Figure 6 is a schematic diagram of the structure of a surface normal estimation device of a target object in an exemplary embodiment of the present application;

[0044] Figure 7 is a schematic diagram of the structure of an electronic device provided by an exemplary embodiment of the present application.

[0045] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. Detailed implementation manners

[0046] The following further describes the present application in detail with reference to the accompanying drawings and embodiments. It can be understood that the embodiments described herein are only used to explain the related invention and do not limit the invention. Additionally, it should be noted that for ease of description, only parts related to the relevant invention are shown in the drawings.

[0047] There are many surface normal estimation methods based on deep learning in the prior art. However, most of these methods take a single RGB image as input and estimate the corresponding surface normal map. Due to the inherent shape - reflectance ambiguity of single - image input, they usually have difficulty in dealing with objects with rich details and unclear appearances. Specifically, a single image may be affected by additional shadows that obscure the details of the object, and the ambient lighting conditions during the capture process may not fully reveal the true geometric features of the scene.

[0048] For this reason, in some embodiments of the present application, a method for estimating the surface normal of an object is provided. As Figure 1 shown, the method includes steps S1 - S3.

[0049] S1. Obtain a first image of the target object under flash conditions and a second image of the target object under non - flash conditions;

[0050] S2. Pre - process the first image and the second image;

[0051] S3. Input the pre - processed images into a trained preset network for normal estimation to obtain the normal map of the target object. Among them, as Figure 2 shown, the preset network includes an encoder module, a diffusion denoising module, and a decoder module.

[0052] In a possible implementation manner, when obtaining the first image of the target object under flash conditions and the second image of the target object under non - flash conditions, it can be adopted Figure 3The imaging device shown includes a camera phone with a flash. To accurately obtain image information under different lighting conditions, the relative position between the real-world object, i.e., the target object, and the camera phone can be adjusted first. Then, a first image under flash conditions can be obtained when the flash is turned on, and a second image under non-flash conditions can be obtained when the flash is turned off. By inputting the pre-processed first image (flash image) and the second image (non-flash image) into the encoder module respectively, extracting their respective latent features, and merging them, the image information of flash and non-flash can be effectively fused, and the complementary characteristics of both in object surface normal estimation can be fully exploited. This operation provides richer input information for subsequent normal estimation, helping to improve the accurate capture of object surface details by the algorithm under different lighting conditions. By merging the features of these two images, the network can achieve a balance between detail capture and noise removal, thus improving the quality of the finally estimated normal map and avoiding local distortion or detail loss that may exist in a single image.

[0053] Figure 3 It also includes a 3D scanner for scanning real-world objects. After scanning a three-dimensional mesh, an accurate mesh representation of the target object can be obtained. The data collected by the 3D scanner will be used for normal map alignment to obtain the aligned real object surface normal map, which is used as a reference for the real data (which can be used as a label) in the preset mesh during training. We need to pre-train the preset network. When training the preset network, first, a training data set and a test data set need to be collected. In this application, synthetic data is used as the training data set, and real data is used as the test data set. In one possible implementation, the training method of the preset network includes: obtaining a synthetic data set and a real data set, where the synthetic data set includes computer-generated images with virtual target objects, and the real data set includes images of real objects taken; pre-processing the synthetic data set and the real data set to obtain flash images and non-flash images suitable for input; training the preset network based on the flash images and the non-flash images.

[0054] When obtaining synthetic data specifically, it is obtained through the blender tool. The blender tool is used to render the real normal map of the object surface under flash and non-flash conditions as synthetic training data. More specifically, the cycles renderer in blender can be used to render and generate synthetic data. Blender is a powerful open source 3D computer graphics software, in which the cycles rendering engine supports real-time rendering. In this application, blender is used to synthesize training data. By placing the photographed objects in three-dimensional space and adding ambient light sources, the camera's intrinsic and extrinsic parameters are set to simulate the results of taking images with an RGB camera in a real scene. At the same time, the surface normal information of the object is also obtained through this software. Refer to Figure 3 , images are taken by a Canon camera as real test data, and the real object is scanned by a 3D scanner to obtain its three-dimensional mesh digital representation, which can then be manually aligned to obtain the surface normal map of the real object. Figure 4 is a schematic diagram of a real data set captured in an exemplary embodiment of the present application, such as Figure 4 As shown, the present application photographed and aligned dozens of objects. This test set can be used as a basic test set for testing normal estimation algorithms, which can better evaluate the quality of normal estimation algorithms.

[0055] In a possible implementation, the preprocessing of the first image and the second image includes: obtaining mask images of the first image and the second image; intercepting the target object based on the mask image to obtain an image corresponding to the main part of the target object; calculating and generating a bounding box based on the image corresponding to the main part to define the area of ​​the main part, for example, calculating the bounding box of the object in four directions in the image through the non-zero pixel value coordinates of the object boundary position in the mask image, thereby providing the object area for further image processing; and using the images defined by the bounding boxes as the preprocessed first image and the preprocessed second image.

[0056] It should be noted that in the training phase, for synthetic data, the mask image of the object can be obtained from the blender renderer. After obtaining the mask image, its pixel value range is divided by 255 to obtain a mask image with pixel values ​​of 0 or 1. This image is multiplied by the RGB image to obtain an RGB image of only the main body of the object. Furthermore, the bounding box can be obtained according to the coordinates of the non-zero position of the boundary of the object in the mask image, and then the bounding box is used to process the output image of the previous step to obtain the synthetic data after the main body of the object is enlarged. This data is used as input in the process of training the preset network. However, since there is no mask in the real data, the RGB image can be cut out to obtain the mask image of the real data.

[0057] In a possible implementation, the preprocessed image is input into a trained preset network for normal estimation to obtain the normal map of the target object, including: inputting the preprocessed first image and the preprocessed second image into the encoder module respectively to extract latent features, obtaining a first feature map corresponding to the first image and a second feature map corresponding to the second image; merging the first feature map, the second feature map and the Gaussian noise map in the channel dimension to obtain a merged feature map; inputting the merged feature map into the diffusion denoising module, and the diffusion denoising module gradually removes the noise in the merged feature map to obtain a normal map feature with object surface details; inputting the normal map feature into the decoder module, and the decoder module restores the normal map feature to the normal map of the target object through deconvolution operation.

[0058] It should be noted that merging the first feature map, the second feature map and the Gaussian noise map in the channel dimension is to let the first feature map and the second feature map guide the denoising process of the Gaussian noise map, so that when the model performs the normal map estimation task, it can simultaneously learn to identify and remove noise. Therefore, the use of the Gaussian noise map not only simulates the noise situations that may be encountered in the real world, but more importantly, it provides a specific guideline or path for the model to follow when dealing with these noises.

[0059] Preferably, referring to Figure 5 , the encoder module is based on the variational autoencoder (VAE) structure and can compress the input image into a low-dimensional latent space. As Figure 5 shown, the preprocessed first image (flash image, Figure 5 C in (f) ) and the second image (non-flash image, Figure 5 C in (nf) ) are input into the variational autoencoder respectively, and feature extraction is performed on the two images respectively to obtain corresponding latent feature maps, and then the two are merged in the channel dimension, which can be represented by formula (1):

[0060] z (fnf) =concat(εC (f) ),ε(C (nf) )) Formula (1)

[0061] where concat is channel dimension merging, z (fnf) is the obtained flash non-flash feature vector, that is, the merged feature map, (εC (f) ),ε(C (nf) )) represents the first feature map and the second feature map. After obtaining z (fnf)After that, Gaussian noise is merged with it to obtain the input data input into the diffusion denoising module. Preferably, the diffusion denoising module is based on a U-shaped network structure, and the merged feature map is input Figure 5 into the denoising U-shaped network shown in the figure, and the noise in the merged feature map is gradually removed to obtain a normal map feature with object surface details. The denoising U-shaped network here is trained. Using the details of the diffusion prior, it can gradually remove the noise in the feature map and restore the normal map feature with object surface details. It can be understood that after the diffusion denoising module inputs the merged feature map, it can gradually remove the noise in the image, and this process effectively restores the details of the object surface. In particular, through gradual denoising, the diffusion denoising module not only retains the detailed information of the surface structure, but also improves the accuracy of the normal map by removing redundant noise parts. This step greatly enhances the detail representation ability of the image, making the normal map estimated from flash and non-flash images more refined, clear, and effectively reducing the impact of reflection differences caused by different light sources on normal estimation.

[0062] It should be noted that when training the diffusion denoising module, Gaussian noise is added to the real normal map to simulate different degrees of noise interference, so that the diffusion denoising module learns to restore the real normal map from the noise. It can be understood that when training the diffusion denoising module, by adding Gaussian noise to the real normal map, different degrees of noise interference are simulated, prompting the module to learn to restore the real normal map from the noise. In this way, the diffusion denoising module can obtain the ability to remove noise from the noise, so that it can still maintain high accuracy when facing complex noise sources (such as noise or blurring that may occur during image shooting). This training method significantly enhances the noise resistance ability of the algorithm, thereby improving the accuracy of object surface normal estimation. Especially in practical applications, noise is often inevitable, so this training strategy provides stronger adaptability for the application of the algorithm in the real environment.

[0063] For example, when training the denoising U-shaped network, in this application, Gaussian noise is added to the normal map of a real object to obtain a real latent space normal feature map with noise where t represents the current time step, y represents the normal map that will be predicted later, and is merged with z (fnf) to obtain the variable z input into the denoising U-shaped network t , as shown in Equation (2):

[0064]

[0065] After that, the denoising U-shaped network will learn the network parameters through the backpropagation process. Here, we use the mean square error between the predicted noise and the true noise as the loss function. During the training process, only the parameters in the denoising U-shaped network are updated. The encoder module uses pre-trained weights and does not update the alignment parameters during training. The time steps are sampled from 1 to T.

[0066] During the inference process, by combining z (fnf) with Gaussian noise ∈, the variable input to the U-shaped denoising network is obtained, as shown in Equations (3) and (4):

[0067]

[0068] where is the predicted normal map at time step T. The arrow above z indicates that this variable is predicted, and the superscript y in the upper right corner indicates that this variable is a normal map. z t represents the variable currently input to the U-shaped network at time step t. The denoising process is shown in Equation (5):

[0069]

[0070] where v θ (z t , t) is the rate value predicted by the denoising U-shaped network for the input z t and time step t. a and b are different weight coefficients. After multiple iterations of denoising, as shown in Equation (6):

[0071]

[0072] After inputting z t into the denoising U-shaped network, the predicted noise at time step t can be obtained. First, calculate where the subscript t->0 at the lower right indicates that after the time step changes from t to 0, by adding the calculated noise back to to obtain the latent space variable at time step t - 1. After performing multi-step denoising, that is, repeating the denoising T times, we will obtain a clean predicted feature map where the subscript 1->0 at the lower right indicates that the time step changes from 1 to 0, and then we send it into the decoder module for decoding.

[0073] It should be noted that there is also a CLIP model in the diffusion denoising module. CLIP (Contrastive Language–Image Pretraining) is a multimodal model proposed by OpenAI that can map text and images into the same embedding space. The core idea of CLIP is to train the model through contrastive learning so that it can understand the semantic relationship between text and images. Specifically, the CLIP model is integrated into the denoising U-net and assists the normal estimation task by inputting text prompts (such as Figure 5 "object geometry" in

[0074] ), that is, providing a high-level semantic information in the denoising U-net to guide the denoising process. Specifically, the text prompt is input into the CLIP model to generate a text embedding vector, which is combined with the input feature map of the denoising U-net as additional context information to help the network better understand the geometric characteristics of the target object. In the normal estimation task, the shape-reflectance ambiguity is a common problem, that is, the network may have difficulty distinguishing the geometric shape of the object surface and the material reflection characteristics. The CLIP model provides clear geometric information through text prompts, guiding the denoising network to pay more attention to the geometric characteristics of the object rather than the influence of materials or lighting, thus effectively alleviating this problem. In a further implementation, the normal map feature is input into the decoder module, and the decoder module restores the normal map feature to the normal map of the target object through deconvolution operations. The decoder module ( Figure 5 the decoder in the variational autoencoder in ) outputs the normal map, which is the finally predicted normal map of the target object, containing rich surface details of the object, and effectively reduces the shape-reflectance ambiguity through diffusion denoising operations. In addition, after obtaining the normal map of the target object, three-dimensional reconstruction is performed based on the normal map of the target object, which can realize the restoration from the normal map to the three-dimensional object form. This process provides a solid foundation for the accurate modeling of the surface details of the object. After obtaining high-quality normal maps, the accurate reconstruction of the object surface shape can be carried out based on these normal maps, and then accurate three-dimensional information can be provided for applications such as three-dimensional modeling, virtual reality, and augmented reality. Especially when estimating normals using flash and non-flash images, high-precision normal maps can provide more realistic and delicate surface details of the object in three-dimensional reconstruction, significantly improving the accuracy and effect of three-dimensional reconstruction.

[0075] In other embodiments of the present application, a device for estimating the surface normal of a target object is provided, as shown in Figure 6 , the device includes:

[0076] An acquisition module 601, configured to acquire a first image of a target object under flash conditions and a second image of the target object under non-flash conditions;

[0077] A processing module 602, configured to preprocess the first image and the second image;

[0078] An estimation module 603, configured to input the preprocessed image into a trained preset network for normal estimation to obtain a normal map of the target object;

[0079] Wherein, the preset network includes an encoder module, a diffusion denoising module, and a decoder module.

[0080] The surface normal estimation device of the target object further includes an evaluation module. The evaluation module calculates the mean angular error (MAE) between the normal vector and the true normal vector according to the normal map output by the preset network, and uses it as a quantitative evaluation index for normal estimation to evaluate the accuracy of normal estimation. By combining flash and non-flash images and using a denoising method with diffusion prior, the surface normal of the target object can be accurately estimated under complex lighting conditions.

[0081] Please refer to the following Figure 7 , which shows a schematic diagram of an electronic device provided by some embodiments of the present application. As Figure 7 shown, the electronic device 2 includes: a processor 200, a memory 201, a bus 202, and a communication interface 203. The processor 200, the communication interface 203, and the memory 201 are connected through the bus 202; a computer program that can run on the processor 200 is stored in the memory 201. When the processor 200 runs the computer program, it executes the surface normal estimation method of the target object provided by any of the foregoing embodiments of the present application. The method includes: acquiring a first image of the target object under flash conditions and a second image of the target object under non-flash conditions; preprocessing the first image and the second image; inputting the preprocessed image into a trained preset network for normal estimation to obtain a normal map of the target object; wherein, the preset network includes an encoder module, a diffusion denoising module, and a decoder module.

[0082] Wherein, the memory 201 may include a high-speed random access memory (RAM: Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 203 (which can be wired or wireless), a communication connection is established between this system network element and at least one other network element, and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.

[0083] The bus 202 can be an ISA bus, a PCI bus, an EISA bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. Among them, the memory 201 is used to store a program. After receiving an execution instruction, the processor 200 executes the program. Any implementation manner of the control method disclosed in any implementation manner of the embodiments of the present application can be applied to or implemented by the processor 200.

[0084] The processor 200 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in hardware or the instructions in software form in the processor 200. The above-mentioned processor 200 can be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 201, and the processor 200 reads the information in the memory 201 and combines its hardware to complete the steps of the method for estimating the surface normal of the target object.

[0085] The embodiment of the present application also provides a computer-readable storage medium corresponding to the method for estimating the surface normal of a target object provided in the foregoing embodiment. A computer program is stored thereon. When the computer program is run by a processor, it will execute the method for estimating the surface normal of a target object provided in any foregoing embodiment.

[0086] In addition, examples of the computer-readable storage medium may further include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other optical or magnetic storage media, which will not be elaborated herein one by one.

[0087] In addition, an embodiment of the present application also provides a computer program product, including a computer program, which when executed by a processor, implements the method for estimating the surface normal of a target object provided in any of the foregoing embodiments. The method includes: obtaining a first image of the target object under flash conditions and a second image of the target object under non-flash conditions; preprocessing the first image and the second image; inputting the preprocessed images into a trained preset network for normal estimation to obtain a normal map of the target object; wherein the preset network includes an encoder module, a diffusion denoising module, and a decoder module.

[0088] Those skilled in the art can understand that each component embodiment of the present application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components in the virtual machine creation device according to the embodiments of the present application.

[0089] As described above, the foregoing are only specific embodiments of the present application that are preferred. However, the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for estimating a surface normal of a target object, characterized in that: The method comprises: Acquire a first image of the target object under a flash condition and a second image under a non-flash condition; Preprocessing the first image and the second image; Inputting the preprocessed image into a trained preset network for normal estimation to obtain a normal map of the target object; The preset network includes an encoder module, a diffusion denoising module and a decoder module.

2. The surface normal estimation method of a target object according to claim 1, characterized in that: The preprocessing of the first image and the second image comprises: Acquire mask images of the first image and the second image; Based on the mask image, the target object is intercepted to obtain an image corresponding to the main part of the target object; A bounding box is calculated and generated based on the image corresponding to the main part to define the area of ​​the main part, and the image defined by the bounding box is used as the preprocessed first image and the preprocessed second image.

3. The surface normal estimation method of a target object according to claim 1, characterized in that: The step of inputting the preprocessed image into a trained preset network to perform normal estimation to obtain a normal map of the target object includes: Inputting the preprocessed first image and the preprocessed second image into the encoder module to extract potential features, respectively, to obtain a first feature map corresponding to the first image and a second feature map corresponding to the second image; Merging the first feature map, the second feature map, and the Gaussian noise map in the channel dimension to obtain a merged feature map; The merged feature map is sequentially input into the diffusion denoising module and the decoder module to obtain a normal map of the target object.

4. The surface normal estimation method of a target object according to claim 1, characterized in that: The step of sequentially inputting the merged feature map into the diffusion denoising module and the decoder module to obtain a normal map of the target object comprises: The merged feature map is input into the diffusion denoising module, and the diffusion denoising module gradually removes the noise in the merged feature map to obtain a normal map feature with surface details of the object; The normal map features are input into the decoder module, and the decoder module restores the normal map features to the normal map of the target object through a deconvolution operation.

5. The surface normal estimation method of a target object according to claim 1, characterized in that: The training method of the preset network includes: Acquire a synthetic data set and a real data set, wherein the synthetic data set includes computer-generated images with virtual target objects, and the real data set includes captured images of real objects; Preprocessing the synthetic data set and the real data set to obtain flash images and non-flash images suitable for input; The preset network is trained based on the flash image and the non-flash image.

6. The surface normal estimation method of a target object according to claim 5, characterized in that: The training method of the preset network also includes: When the diffusion denoising module is trained, Gaussian noise is added to the real normal map to simulate noise interference of different degrees, so that the diffusion denoising module learns to restore the real normal map from the noise.

7. The surface normal estimation method of a target object according to claim 1, characterized in that: After obtaining the normal map of the target object, the method further includes: performing three-dimensional reconstruction based on the normal map of the target object.

8. A surface normal estimation device for a target object, characterized in that: The device comprises: An acquisition module, used for acquiring a first image of a target object under a flash condition and a second image under a non-flash condition; A processing module, used for preprocessing the first image and the second image; An estimation module, used for inputting the preprocessed image into a trained preset network to perform normal estimation, so as to obtain a normal map of the target object; The preset network includes an encoder module, a diffusion denoising module and a decoder module.

9. An electronic device, comprising a memory and a processor, characterized in that: The memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor executes the surface normal estimation method of the target object as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the surface normal estimation method of the target object as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Surface normal estimation method and system for transparent object

    CN114972464A

  • Surface defect detection method and system based on multi-light-source cooperation

    CN115272175A

  • Abrasion surface topography luminosity three-dimensional reconstruction method and system fused with all-light-source image

    CN116385520A

  • Image enhancement method and device based on diffusion model, equipment and storage medium

    CN116664450A

  • Generation method and device of stylized image generation model, equipment and storage medium

    CN116740204A