3D Reconstruction Method and Device for Target Object, Electronic Device, Storage Medium

By training the Gaussian model and the dark light enhancement network, and using heat maps and RGB images for joint training, the problem of insufficient three-dimensional reconstruction accuracy under low light conditions is solved, and a higher precision three-dimensional reconstruction is achieved.

CN119832169BActive Publication Date: 2025-07-11BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510315199.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-11
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

The existing 3D Gaussian speckle method increases image noise and reduces contrast under low light conditions, resulting in blurred object contours and inability to obtain effective information, limiting the accuracy of three-dimensional reconstruction.

Method used

By training Gaussian models and dark light enhancement networks, combined training is used to generate higher quality three-dimensional reconstruction results.

Benefits of technology

Improve the three-dimensional reconstruction accuracy of the target object in low-light environments to generate higher quality mesh results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119832169B_ABST
    Figure CN119832169B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method and apparatus for three-dimensional reconstruction of a target object, an electronic device, and a storage medium, belonging to the technical field of image processing. The method includes: training a Gaussian model based on a first sample set to obtain a first Gaussian network; the first sample set includes a plurality of first RGB images under low-light conditions; training a neural network model based on a second sample set to obtain a low-light enhancement network; inputting the first RGB image and the first heat map corresponding to the first RGB image into the low-light enhancement network, and inputting the output of the low-light enhancement network and the first heat map into the first Gaussian network for joint training to obtain a second Gaussian network; performing three-dimensional reconstruction of the target object based on the second Gaussian network. The method and apparatus for three-dimensional reconstruction of a target object, the electronic device, and the storage medium provided by the present disclosure can improve the accuracy of three-dimensional reconstruction of the target object under low-light conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure belongs to the technical field of image processing, and more particularly, relates to a method and apparatus for three-dimensional reconstruction of an object, an electronic device, and a storage medium. Background Art

[0002] The technology of scene reconstruction and novel view synthesis based on visible light has achieved remarkable achievements, and its application scope has penetrated into many fields, bringing great convenience to people's work and life. In the field of virtual reality (VR), this technology can create an immersive experience, making users feel as if they are on the scene; in the field of autonomous driving, it helps vehicles achieve accurate environmental perception and path planning to ensure driving safety; in the field of robotics, it helps robots better understand the surrounding environment and achieve efficient interaction with the environment; in the construction industry, it can be used for building design and construction monitoring to improve design quality and construction efficiency; in addition, in the field of digital twins, by constructing a digital copy of the real scene, it provides strong support for various decisions.

[0003] 3D Gaussian speckle (3DGS) can achieve real-time rendering while maintaining high visual fidelity by using a large number of 3D Gaussians to represent the scene, making a qualitative leap in the efficiency of 3D reconstruction and view synthesis. However, most current 3DGS methods mainly process RGB image data under ideal lighting conditions. In actual application scenarios, such as at night, in indoor low-light environments, or in the presence of severe occlusion, RGB images will face many challenges. Due to the photon starvation effect, the performance will significantly decline under low-light conditions, resulting in increased image noise, reduced contrast, making the outline of the object blurred, and even unable to obtain sufficient effective image information, which greatly limits the application effect and reconstruction accuracy of 3DGS technology in these special scenarios. Summary of the Invention

[0004] The purpose of the present disclosure is to provide a method and apparatus for three-dimensional reconstruction of an object, an electronic device, and a storage medium to improve the accuracy of three-dimensional reconstruction of an object under low-light conditions.

[0005] In the first aspect of the embodiments of the present disclosure, a method for three-dimensional reconstruction of an object is provided, including:

[0006] Training a Gaussian model based on a first sample set to obtain a first Gaussian network; the first sample set includes a plurality of first RGB images under low-light conditions;

[0007] Train a neural network model based on a second sample set to obtain a low-light enhancement network; the second sample set includes multiple groups of image pairs, and any group of image pairs includes a second RGB image of a first object under low-light conditions, a third RGB image of the first object under ideal lighting conditions, and a heat map of the first object;

[0008] Input the first RGB image and the first heat map corresponding to the first RGB image into the low-light enhancement network, and input the output of the low-light enhancement network and the first heat map into the first Gaussian network for joint training to obtain a second Gaussian network;

[0009] Perform three-dimensional reconstruction on the target object based on the second Gaussian network.

[0010] In a second aspect of the embodiments of the present disclosure, there is provided a three-dimensional reconstruction device for a target object, including:

[0011] A first training module, configured to train a Gaussian model based on a first sample set to obtain a first Gaussian network; the first sample set includes multiple first RGB images under low-light conditions;

[0012] A second training module, configured to train a neural network model based on a second sample set to obtain a low-light enhancement network; the second sample set includes multiple groups of image pairs, and any group of image pairs includes a second RGB image of a first object under low-light conditions, a third RGB image of the first object under ideal lighting conditions, and a heat map of the first object;

[0013] A joint training module, configured to input the first RGB image and the first heat map corresponding to the first RGB image into the low-light enhancement network, and input the output of the low-light enhancement network and the first heat map into the first Gaussian network for joint training to obtain a second Gaussian network;

[0014] A three-dimensional reconstruction module, configured to perform three-dimensional reconstruction on the target object based on the second Gaussian network.

[0015] In a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, and when the processor executes the computer program, the steps of the above-mentioned three-dimensional reconstruction method for a target object are implemented.

[0016] In a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned three-dimensional reconstruction method for a target object are implemented.

[0017] The beneficial effects of the three-dimensional reconstruction method, device, electronic device, and storage medium for a target object provided by the embodiments of the present disclosure are as follows:

[0018] In the embodiments of the present disclosure, a Gaussian model is trained using a heat map and an RGB image to perform three-dimensional reconstruction on a target object. The characteristic that thermal imaging technology is not affected by lighting conditions can be fully utilized to help identify the contour information of objects in a low-light environment, thereby improving the accuracy of three-dimensional reconstruction in a low-light environment.

[0019] Among them, the training process is divided into a pre-training stage and an end-to-end training stage. In the pre-training stage, the Gaussian model is pre-trained to obtain a first Gaussian network, and the neural network model is pre-trained to obtain a low-light enhancement network. The low-light enhancement network can enhance low-light images to make them closer to the image quality under ideal lighting. The above pre-training steps can reduce the convergence time in the subsequent end-to-end training process. In the end-to-end training stage, the low-light enhancement network and the first Gaussian network are jointly trained. In this process, the low-light enhancement network and the first Gaussian network are optimized through various constraint conditions to obtain a second Gaussian network with better performance.

[0020] Therefore, the embodiments of the present disclosure improve the Gaussian model's understanding ability of the scene through the low-light enhancement network and the first heat map, can generate higher-quality mesh results, and thus achieve higher-precision three-dimensional reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0022] Figure 1 It is a schematic flowchart of a method for three-dimensional reconstruction of a target object provided by an embodiment of the present disclosure;

[0023] Figure 2 It is a schematic diagram of the end-to-end training process provided by an embodiment of the present disclosure;

[0024] Figure 3 It is a schematic diagram of the low-light enhancement network provided by an embodiment of the present disclosure;

[0025] Figure 4 It is a structural block diagram of a three-dimensional reconstruction device for a target object provided by an embodiment of the present disclosure;

[0026] Figure 5 It is a schematic block diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] In the following description, specific details such as specific system architectures, technologies, etc. are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should clearly understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present disclosure.

[0028] To make the objectives, technical solutions, and advantages of the present disclosure clearer, the following will be described through specific embodiments in conjunction with the accompanying drawings.

[0029] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a three-dimensional reconstruction method for a target object provided in an embodiment of the present disclosure. The method includes:

[0030] S101: Train a Gaussian model based on a first sample set to obtain a first Gaussian network; the first sample set includes a plurality of first RGB images under low-light conditions.

[0031] In this embodiment, a Gaussian model is trained using heatmaps and RGB images to perform three-dimensional reconstruction of the target object based on the trained Gaussian model. The training process is divided into two stages, namely the pre-training stage and the end-to-end training stage. In the pre-training stage, the first Gaussian network and the low-light enhancement network are pre-trained to reduce the convergence time of the subsequent end-to-end training process; in the end-to-end training stage, the pre-trained first Gaussian network is fine-tuned using heatmaps and RGB images that have passed through the low-light enhancement network to obtain a second Gaussian network.

[0032] Specifically, a plurality of first RGB images under low-light conditions can be collected in advance as the first sample set, and a Gaussian model is trained based on the first sample set, enabling the Gaussian model to learn the characteristics and distribution laws of RGB images under low-light conditions. For example, the Gaussian model can learn the approximate shape and contour information of objects under low light, and this information can be used as the basic information for constructing a mesh. Among them, a mesh is a geometric structure for discretely representing an object or space, which consists of a series of vertices, edges, and faces and is used to accurately describe the shape and surface of an object. A high-quality mesh can accurately present the details and complex geometric features of an object and is the basic information for three-dimensional modeling.

[0033] Considering that it is not in line with practical applications to achieve the effect of converting a low-light image to an ideal illumination image when the ideal illumination image is known, in the end-to-end training stage, this embodiment only uses the heat map and the low-light RGB image, without using the ideal illumination image. Correspondingly, in the pre-training stage, a Gaussian model is trained using a first sample set composed of multiple low-light RGB images. Although there will be problems such as missing texture details and inaccurate depth estimation during the training process, initialization data can be established based on partial observable information, and it has better view consistency compared to the low-light enhanced RGB images in the end-to-end training stage.

[0034] S102: Train a neural network model based on the second sample set to obtain a low-light enhancement network; the second sample set includes multiple groups of image pairs, and any group of image pairs includes a second RGB image of the first object under low-light conditions, a third RGB image of the first object under ideal illumination conditions, and a heat map of the first object.

[0035] In this embodiment, the second RGB image of the same object (such as the first object) under low-light conditions, the third RGB image of the same object under ideal illumination conditions, and the heat map can be collected in advance. The above images corresponding to the same object are used as a group of image pairs, and multiple groups of image pairs corresponding to multiple objects are collected in the same way to obtain the second sample training set.

[0036] Training a neural network model based on the second sample set can obtain a low-light enhancement network. The neural network model here can be implemented using a Retinex-Net, a Latent Disentangled Enhancement Network (LDE-Net), a Lightweight YUVTransformer Network (LYT-Net), etc.

[0037] During the training process, taking advantage of the fact that the heat map can obtain the object contour under low-light conditions, it is incorporated into the low-light enhancement network to improve the low-light enhancement effect and make the output image of the low-light enhancement network approach the third RGB image.

[0038] S103: Input the first RGB image and the first heat map corresponding to the first RGB image into the low-light enhancement network, and input the output of the low-light enhancement network and the first heat map into the first Gaussian network for joint training to obtain the second Gaussian network.

[0039] Please refer to Figure 2 , Figure 2Schematic diagram of the end-to-end training process provided by an embodiment of the present disclosure. In the pre-training stage, we have pre-trained the first Gaussian network using the first RGB image in advance. However, the first Gaussian network is a Gaussian model that takes a low-light RGB image as input and outputs a low-light RGB image, while our goal is to be able to input a low-light RGB image and a heat map and output a low-light enhanced image and a heat map. Therefore, in the end-to-end training stage, the first RGB image and the first heat map corresponding to the first RGB image are input into the low-light enhancement network, and the output of the low-light enhancement network and the first heat map are input into the first Gaussian network for joint training. The low-light enhancement network is used to make the RGB information input into the first Gaussian network more, and at the same time, combined with the information of the first heat map, the three-dimensional Gaussian sphere generated by the first Gaussian network is adjusted.

[0040] In the end-to-end training stage, the understanding ability of the Gaussian model for the scene is improved through the low-light enhancement network and the first heat map, and a higher-quality mesh result can be generated, thereby achieving higher-precision three-dimensional reconstruction.

[0041] S104: Perform three-dimensional reconstruction on the target object based on the second Gaussian network.

[0042] In this embodiment, through end-to-end joint training, the second Gaussian network has learned richer and more accurate image feature information, including the features after low-light image enhancement and the auxiliary information provided by the heat map. Using this information, the second Gaussian network can more accurately obtain information such as the shape, position, and texture of the target object, thereby performing three-dimensional reconstruction on the target object.

[0043] It can be concluded from the above that this embodiment uses a heat map and an RGB image to train a Gaussian model to perform three-dimensional reconstruction on a target object based on the Gaussian model, which can make full use of the characteristic that thermal imaging technology is not affected by lighting conditions to help identify the contour information of objects in a low-light environment, thereby improving the accuracy of three-dimensional reconstruction in a low-light environment.

[0044] Among them, the training process is divided into a pre-training stage and an end-to-end training stage. In the pre-training stage, the Gaussian model is pre-trained to obtain the first Gaussian network, and the neural network model is pre-trained to obtain the low-light enhancement network. The low-light enhancement network can enhance the low-light image to make it closer to the image quality under ideal lighting. The above pre-training steps can reduce the convergence time of the subsequent end-to-end training process. In the end-to-end training stage, the low-light enhancement network and the first Gaussian network are jointly trained. In this process, the low-light enhancement network and the first Gaussian network are optimized through various constraint conditions to obtain a second Gaussian network with better performance.

[0045] Therefore, in this embodiment, the low-light enhancement network and the first heat map are used to improve the Gaussian model's understanding ability of the scene, enabling the generation of higher-quality mesh results, thereby achieving more accurate 3D reconstruction.

[0046] In an embodiment of the present disclosure, the first RGB image is different from the second RGB image.

[0047] In this embodiment, a supervised learning method is used to train the low-light enhancement network, that is, the third RGB image under ideal illumination conditions is used as the ground truth after enhancing the second RGB image under low-light conditions to supervise the low-light enhancement effect.

[0048] Considering that in the end-to-end training stage, the goal of this embodiment is to do a self-supervised task, that is, to achieve high-quality enhancement of low-light images without using ideal illumination images to supervise low-light images. Therefore, in this stage, this embodiment does not use ideal illumination images. In the pre-training stage, only to pre-train a low-light enhancement network, so paired second RGB images under low-light conditions and third RGB images under ideal illumination conditions are used for constraint.

[0049] It can be concluded from the above that in the pre-training stage, this embodiment uses different scenarios (i.e., different RGB images) from the end-to-end training stage for supervised training, so as to obtain a better low-light enhancement network.

[0050] In an embodiment of the present disclosure, training the neural network model based on the second sample set includes:

[0051] Performing parameter adjustment operations multiple times until the stop condition is met;

[0052] The parameter adjustment operation includes:

[0053] Enhance the second RGB image of the first object under low-light conditions to obtain a first enhanced image;

[0054] Fuse the heat map of the first object and the first enhanced image to obtain a second enhanced image;

[0055] Determine a first loss function based on the second enhanced image and the third RGB image of the first object under ideal illumination conditions, and adjust the parameters of the neural network model based on the first loss function;

[0056] The stop condition is: the first loss function is less than the first threshold, or the number of iterations reaches the set number of times.

[0057] Please refer to Figure 3 , Figure 3Schematic diagram of a low-light enhancement network provided by an embodiment of the present disclosure. In this embodiment, first, a neural network model such as a Retinex-Net, a Latent Disentangled Enhancement Network (LDE-Net), or a Lightweight YUV Transformer Network (LYT-Net) can be used to enhance a second RGB image under low-light conditions to obtain a first enhanced image; then, a heat map is incorporated into the first enhanced image to improve the low-light enhancement effect, so that the obtained second enhanced image approaches a third RGB image under ideal illumination conditions.

[0058] Among them, the fusion of the heat map and the first enhanced image can be implemented by a network such as a Generative Adversarial Network (GAN) or a Variational Autoencoder (VAE).

[0059] Specifically, in the training process, the parameters of the neural network model are continuously iteratively optimized by repeatedly performing parameter adjustment operations to achieve a better image enhancement effect. In each iteration, a first loss function is calculated based on the difference between the output of the neural network model (i.e., the second enhanced image) and the ground truth (i.e., the third RGB image under ideal illumination conditions). The iteration can stop when the first loss function is less than a pre-set first threshold, or when the number of iterations reaches a set number.

[0060] It can be concluded from the above that in this example, by continuously iteratively adjusting the parameters of the neural network model and comprehensively using the information of the low-light RGB image and the heat map, the image enhancement effect is gradually optimized, enabling the model to generate high-quality images close to the ideal illumination conditions when facing new low-light images.

[0061] In an embodiment of the present disclosure, enhancing a second RGB image of a first object under low-light conditions to obtain a first enhanced image includes:

[0062] Decomposing the second RGB image of the first object under low-light conditions into a reflection map and an illumination map;

[0063] Enhancing the reflection map and the illumination map respectively;

[0064] Combining the enhanced reflection map and the enhanced illumination map into a first enhanced image.

[0065] In this embodiment, a specific implementation manner for enhancing the second RGB image under low-light conditions is given. Based on the Retinex-Net, the second RGB image can be decomposed into a reflection map and an illumination map, and then the reflection map and the illumination map from the second RGB image are enhanced respectively, and then the enhanced images are combined into a first enhanced image.

[0066] Among them, enhancing the decomposed reflection map can highlight the details such as the texture and edges of the object, and enhancing the decomposed illumination map can adjust the problem of uneven illumination, making the details in the dark areas also visible.

[0067] As can be seen from the above, in this embodiment, the second RGB image is decomposed into a reflection map and an illumination map, and then the reflection map and the illumination map are enhanced respectively, so that the enhanced image can contain more details and more accurate illumination information.

[0068] On this basis, the first loss function of the low-light enhancement network includes a reconstruction loss, a reflection invariance loss, an illumination smoothness loss, a self-feature retention loss, and an enhancement loss. The calculation processes of each loss function are as follows:

[0069] (1) Reconstruction loss: According to the Retinex theory, the decomposed reflection map R and illumination map I can be recombined into the original image , which ensures the integrity of information. Therefore, the reconstruction loss is calculated as follows:

[0070]

[0071] In the formula, is the reflection map of the low-light image, is the reflection map of the ideal illumination image, is the illumination map of the low-light image, is the illumination map of the ideal illumination image, is the low-light image (i.e., the second RGB image), is the ideal illumination image (i.e., the third RGB image).

[0072] (2) Reflection invariance loss: According to the Retinex theory, the reflectance reflects the inherent characteristics of the object itself and is not affected by illumination and imaging devices. That is to say, the reflection maps obtained by decomposing the images captured under different illumination conditions in the same scene should be as consistent as possible. That is, the reflection map of the low-light image decomposed by the network should be as consistent as possible with the reflection map of the ideal illumination image and the reflection map of the low-light enhanced image. Accordingly, the reflection invariance loss is calculated as follows:

[0073]

[0074] In the formula, the part regarding the reflection map of the ideal illumination image is only used in the pre-training stage, because in the end-to-end training, we do not use the ideal illumination image for supervision. The low-light enhanced image used in the pre-training stage represents the first enhanced image obtained through the low-light enhancement network in the end-to-end stage.

[0075] (3) Illumination smoothness loss: To better guide the network to decompose the illumination map, we use the loss in the Darkness-free infrared and visible image fusion (DIVFusion) method and the loss :

[0076]

[0077] where ∇ is the Sobel operator, including the x and y directions, represents the exponential calculation with the natural logarithm e as the base; is a small positive constant to prevent division by zero; is a parameter that controls the shape of the mutual consistency loss and aims to increase the medium-gradient part of the image.

[0078] (4) Self-feature retention loss: To ensure better constraints on end-to-end training when there is no ideal illumination image, we also use the self-feature retention loss of the EnlightenGAN (Deep light enhancement generative adversarial network without paired supervision) to constrain the input low-light image and the output low-light enhanced image in terms of the VGG feature distance.

[0079]

[0080] where ϕ represents the feature map extracted by the pre-trained VGG19 network, i represents the i-th max pooling layer, j represents the j-th convolutional layer after the i-th max pooling layer, i = 5, j = 1; and are the sizes of the extracted feature maps.

[0081] (4) Enhancement loss: For the output second enhanced image (i.e., the final enhanced image Figure 3 in ) and the real ideal illumination image, we calculate the enhancement loss, which is only used in the pre-training stage because in end-to-end training, we do not use the ideal illumination image for supervision. The enhancement loss includes the perceptual loss , loss. The perceptual loss estimates the output result through the pre-trained VGG19 network ​ and the ground truth between the similarities.

[0082]

[0083] In the formula, and represent weighted hyperparameters, represents the L2 norm.

[0084] In an embodiment of the present disclosure, the three-dimensional reconstruction method of the target object further includes:

[0085] Performing super-resolution operation on the second heat map output by the second Gaussian network to obtain a third heat map;

[0086] Performing super-resolution operation on the fourth RGB image output by the second Gaussian network to obtain a fifth RGB image.

[0087] In this embodiment, based on the RGB image and heat map of the target object, the second Gaussian network can generate a second heat map and a fourth RGB from a new perspective. Further, performing super-resolution operation on the second heat map can obtain a high-resolution third heat map, and performing super-resolution operation on the fourth RGB image can obtain a high-resolution fifth RGB image, providing a high-quality data basis for subsequent image analysis and processing.

[0088] Specifically, the super-resolution operations on the second heat map and the fourth RGB image can both be implemented by existing super-resolution networks. In the end-to-end training stage, the low-light enhancement network, the first Gaussian network, and the two super-resolution networks are jointly trained to form an overall network architecture. In this process, through various constraint conditions, each part is optimized.

[0089] Among them, connecting the low-light enhancement network in front of the first Gaussian network is to enable the RGB image input to the first Gaussian network to have more texture details, which can better help the construction of depth information to generate a better mesh effect. Connecting the super-resolution network behind the first Gaussian network is to reduce the memory consumption required for training the first Gaussian network and increase the training speed to achieve a more economical and feasible training process, avoiding unnecessary consumption caused by inputting high-resolution images into the first Gaussian network. And the first Gaussian network can impose a perspective consistency constraint on the images generated by the low-light enhancement network and the super-resolution network, making the images in a three-dimensional scene rather than discrete, so as to avoid problems such as artifacts generated by new perspective generation and deterioration of mesh quality.

[0090] As can be seen from the above, in this embodiment, by jointly training the low-light enhancement network, the first Gaussian network, and two super-resolution networks, an overall network architecture is formed, which can realize obtaining high-resolution images and high-resolution heatmaps from new perspectives using heatmaps and RGB images in low-light scenarios, and then obtaining high-quality mesh results.

[0091] In an embodiment of the present disclosure, the three-dimensional reconstruction method of the target object further includes:

[0092] Guiding the super-resolution operation of the second heatmap based on the fourth RGB image;

[0093] Guiding the super-resolution operation of the fourth RGB image based on the second heatmap.

[0094] In this embodiment, the RGB image has rich visual detail information such as texture and color. By guiding the super-resolution operation of the second heatmap with the fourth RGB image, it can provide more accurate structural and edge information guidance for heatmap super-resolution; at the same time, the second heatmap can clearly identify the contour information of the object in the low-light environment. By guiding the super-resolution operation of the fourth RGB image with the second heatmap, it can provide auxiliary information to supplement details for the super-resolution operation of the fourth RGB image.

[0095] As can be seen from the above, in the process of the super-resolution operation of the second heatmap and the super-resolution operation of the fourth RGB image in this embodiment, another modality is used to complement the details that the original modality fails to capture, which is beneficial to improving the accuracy of the super-resolution operation.

[0096] In an embodiment of the present disclosure, the loss function of the first Gaussian network is specifically:

[0097]

[0098] Among them, represents the loss function of the first Gaussian network, represents the inverse tone curve, represents the pixel value of the ground truth image, represents the pixel value of the output image, represents the L1 norm, represents the hyperparameter.

[0099] In this embodiment, considering that when calculating the loss function of the Gaussian model based on the existing MSE loss function or L1 loss function, the pixels in the dark tend to have too small weights, while the weights of the bright parts are relatively large, which will cause some deviations in the training of low-light images. Therefore, before calculating the loss function of the first Gaussian network in this embodiment, the image is first passed through an inverse tone curve to rebalance the pixel weights, and finally the improved loss function is obtained. Among them, .

[0100] As can be seen from the above, in this embodiment, the pixel weights are rebalanced through the inverse tone curve, avoiding the problem of too small weights for dark pixels, enabling the model to make more full use of the information in the dark areas during training, reducing the deviation in the training of dark images, and thus more accurately learning the features and patterns in the dark images, improving the modeling ability for dark scenes.

[0101] Corresponding to the three-dimensional reconstruction method of the target object in the above embodiment, Figure 4 is a structural block diagram of a three-dimensional reconstruction device for a target object provided by an embodiment of the present disclosure. For ease of description, only parts related to the embodiments of the present disclosure are shown. Refer to Figure 4 The three-dimensional reconstruction device 20 of the target object includes: a first training module 21, a second training module 22, a joint training module 23, and a three-dimensional reconstruction module 24.

[0102] Among them, the first training module 21 is used to train a Gaussian model based on a first sample set to obtain a first Gaussian network; the first sample set includes a plurality of first RGB images under dark light conditions;

[0103] The second training module 22 is used to train a neural network model based on a second sample set to obtain a dark light enhancement network; the second sample set includes multiple groups of image pairs, and any group of image pairs includes a second RGB image of the first object under dark light conditions, a third RGB image of the first object under ideal lighting conditions, and a thermal map of the first object;

[0104] The joint training module 23 is used to input the first RGB image and the first thermal map corresponding to the first RGB image into the dark light enhancement network, and input the output of the dark light enhancement network and the first thermal map into the first Gaussian network for joint training to obtain a second Gaussian network;

[0105] The three-dimensional reconstruction module 24 is used to perform three-dimensional reconstruction on the target object based on the second Gaussian network.

[0106] In an embodiment of the present disclosure, the first RGB image is different from the second RGB image.

[0107] In an embodiment of the present disclosure, the second training module 22 is specifically used for:

[0108] Performing parameter adjustment operations multiple times until the stop condition is met;

[0109] The parameter adjustment operations include:

[0110] Enhancing the second RGB image of the first object under dark light conditions to obtain a first enhanced image;

[0111] Fusing the thermal map of the first object and the first enhanced image to obtain a second enhanced image;

[0112] Determine a first loss function based on a second enhanced image and a third RGB image of the first object under ideal illumination conditions, and adjust the parameters of the neural network model based on the first loss function;

[0113] The stopping condition is: the first loss function is less than a first threshold, or the number of iterations reaches a set number.

[0114] In an embodiment of the present disclosure, the second training module 22 is further specifically configured to:

[0115] Decompose the second RGB image of the first object under low-light conditions into a reflection map and an illumination map;

[0116] Enhance the reflection map and the illumination map respectively;

[0117] Synthesize the enhanced reflection map and the enhanced illumination map into a first enhanced image.

[0118] In an embodiment of the present disclosure, the joint training module 23 is specifically configured to:

[0119] Perform super-resolution operation on the second heat map output by the second Gaussian network to obtain a third heat map;

[0120] Perform super-resolution operation on the fourth RGB image output by the second Gaussian network to obtain a fifth RGB image.

[0121] In an embodiment of the present disclosure, the joint training module 23 is further specifically configured to:

[0122] Guide the super-resolution operation of the second heat map based on the fourth RGB image;

[0123] Guide the super-resolution operation of the fourth RGB image based on the second heat map.

[0124] In an embodiment of the present disclosure, the loss function of the first Gaussian network is specifically:

[0125]

[0126] Wherein, represents the loss function of the first Gaussian network, represents the inverse tone curve, represents the pixel value of the ground truth image, represents the pixel value of the output image, represents the L1 norm, represents a hyperparameter.

[0127] See Figure 5 , Figure 5 is a schematic block diagram of an electronic device provided by an embodiment of the present disclosure. As Figure 5The electronic device 300 in the present embodiment shown may include: one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The above-mentioned processors 301, input devices 302, output devices 303, and memories 304 communicate with each other through a communication bus 305. The memory 304 is used to store a computer program, and the computer program includes program instructions. The processor 301 is used to execute the program instructions stored in the memory 304. Among them, the processor 301 is configured to call the program instructions to execute the functions of each module / unit in the above-mentioned device embodiments, for example Figure 4 the functions of the modules 21 to 24 shown.

[0128] It should be understood that in the embodiments of the present disclosure, the so-called processor 301 may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.

[0129] The input device 302 may include a touchpad, a fingerprint acquisition sensor (for acquiring the fingerprint information and the direction information of the fingerprint of the user), a microphone, etc., and the output device 303 may include a display (such as an LCD), a speaker, etc.

[0130] The memory 304 may include a read-only memory and a random access memory, and provide instructions and data to the processor 301. A part of the memory 304 may also include a non-volatile random access memory. For example, the memory 304 may also store information about the device type.

[0131] In specific implementation, the processors 301, input devices 302, and output devices 303 described in the embodiments of the present disclosure may implement the implementation manners described in the first embodiment and the second embodiment of the three-dimensional reconstruction method of the target object provided by the embodiments of the present disclosure, and may also implement the implementation manner of the electronic device described in the embodiments of the present disclosure, which will not be elaborated herein.

[0132] In another embodiment of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, all or part of the processes in the methods of the above embodiments are implemented. It can also be completed by instructing relevant hardware through the computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0133] The computer-readable storage medium can be an internal storage unit of the electronic device in any of the foregoing embodiments, such as the hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk equipped on the electronic device, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the computer-readable storage medium can also include both the internal storage unit and the external storage device of the electronic device. The computer-readable storage medium is used to store the computer program and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store the data that has been output or will be output.

[0134] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present disclosure.

[0135] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described electronic devices and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0136] In several embodiments provided by this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed coupling or direct coupling or communication connection between each other can be an indirect coupling or communication connection through some interfaces or units, and can also be an electrical, mechanical or other form of connection.

[0137] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present disclosure.

[0138] In addition, in each embodiment of the present disclosure, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0139] The above is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of various equivalent modifications or replacements, and these modifications or replacements should all be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A three-dimensional reconstruction method for a target object, characterized in that, Including: Training a Gaussian model based on a first sample set to obtain a first Gaussian network; The first sample set includes multiple first RGB images under low-light conditions; Training a neural network model based on a second sample set to obtain a low-light enhancement network; the second sample set includes multiple pairs of images, and any pair of images includes a second RGB image of a first object under low-light conditions, a third RGB image of the first object under ideal illumination conditions, and a thermal map of the first object; Inputting the first RGB image and the first thermal map corresponding to the first RGB image into the low-light enhancement network, and inputting the output of the low-light enhancement network and the first thermal map into the first Gaussian network for joint training to obtain a second Gaussian network; Performing three-dimensional reconstruction on a target object based on the second Gaussian network; The training of the neural network model based on the second sample set includes: Performing parameter adjustment operations multiple times until a stop condition is met; The parameter adjustment operation includes: Enhancing the second RGB image of the first object under low-light conditions to obtain a first enhanced image; Fusing the thermal map of the first object and the first enhanced image to obtain a second enhanced image; Determining a first loss function based on the second enhanced image and the third RGB image of the first object under ideal illumination conditions, and adjusting the parameters of the neural network model based on the first loss function; The stop condition is: the first loss function is less than a first threshold, or the number of iterations reaches a set number.

2. The three-dimensional reconstruction method of the target object according to claim 1, wherein, The first RGB image is different from the second RGB image.

3. The three-dimensional reconstruction method of the target object according to claim 1, wherein Enhancing the second RGB image of the first object under low-light conditions to obtain a first enhanced image, including: Decomposing the second RGB image of the first object under low-light conditions into a reflection map and an illumination map; Enhancing the reflection map and the illumination map respectively; Combining the enhanced reflection map and the enhanced illumination map into a first enhanced image.

4. The three-dimensional reconstruction method of the target object according to claim 1, characterized in that, Also including: Performing super-resolution operation on the second thermal map output by the second Gaussian network to obtain a third thermal map; Performing super-resolution operation on the fourth RGB image output by the second Gaussian network to obtain a fifth RGB image.

5. The three-dimensional reconstruction method of the target object according to claim 4, characterized in that, Also including: Guiding the super-resolution operation of the second thermal map based on the fourth RGB image; Guiding the super-resolution operation of the fourth RGB image based on the second thermal map.

6. The three-dimensional reconstruction method of the target object according to claim 1, wherein The loss function of the first Gaussian network is specifically: Among them, represents the loss function of the first Gaussian network, represents the inverse tone curve, represents the pixel value of the ground truth image, represents the pixel value of the output image, represents the L1 norm, represents the hyperparameter.

7. A three-dimensional reconstruction device for a target object, characterized in that, Including: A first training module for training a Gaussian model based on a first sample set to obtain a first Gaussian network; The first sample set includes multiple first RGB images under low-light conditions; A second training module for training a neural network model based on a second sample set to obtain a low-light enhancement network; the second sample set includes multiple pairs of images, and any pair of images includes a second RGB image of a first object under low-light conditions, a third RGB image of the first object under ideal illumination conditions, and a thermal map of the first object; A joint training module, configured to input the first RGB image and the first heat map corresponding to the first RGB image into the low-light enhancement network, and input the output of the low-light enhancement network and the first heat map into the first Gaussian network for joint training to obtain a second Gaussian network; A 3D reconstruction module, configured to perform 3D reconstruction on the target object based on the second Gaussian network; The second training module is specifically configured to: Perform parameter adjustment operations multiple times until a stop condition is met; The parameter adjustment operations include: Enhance the second RGB image of the first object under low-light conditions to obtain a first enhanced image; Fuse the heat map of the first object and the first enhanced image to obtain a second enhanced image; Determine a first loss function based on the second enhanced image and the third RGB image of the first object under ideal lighting conditions, and adjust the parameters of the neural network model based on the first loss function; The stop condition is: the first loss function is less than a first threshold, or the number of iterations reaches a set number.

8. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Point cloud three-dimensional reconstruction method based on RGB data and generative adversarial network

    CN111899328A

  • Dark light image enhancement processing method, device, equipment, system and storage medium

    CN118037575A