Reconstruction model training method, reconstruction method, device and electronic equipment
Patent Information
- Application Number
- CN202510182146.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2026-08-18
AI Technical Summary
尽管传统的MVS重建流程,在室内外场景中都表现良好,但是它们在重建更精细的细节上仍有困难,即表面过于嘈杂甚至不完整
[0022]In this embodiment, firstly, multi-view images of the sample object are acquired and input into the reconstruction model. The reconstruction model performs 3D reconstruction on the multi-view images of the sample object to obtain the corresponding 3D model. The sample object and the 3D model are compared using a target loss function to obtain a target loss value. The reconstruction model is iteratively trained based on the target loss value to obtain an optimized reconstruction model. If the target loss function includes a first loss function for optimizing inconsistencies between multi-view images, that is, the first loss function can optimize the inconsistencies between multi-view images, making the images from different perspectives more consistent in certain features or representations; if the target loss function includes a second loss function for optimizing low-texture surfaces, for some images with less surface texture information, the second loss function can optimize low-texture surfaces, improving the quality or accuracy of these areas in the processing process, thereby improving the quality of reconstruction or rendering; if the target loss function includes a third loss function for optimizing fine-grained structures, for some fine structures on the sample object, the third loss function can enable the model to better capture and reconstruct these details, making the processed image perform better in these detailed structures. By applying at least one of the three different loss functions mentioned above to the target loss function, the multi-view images are processed comprehensively and meticulously to achieve better processing results and more accurate target loss value calculation, thereby improving the model's image rendering and 3D reconstruction performance.
Smart Images

Figure CN122597631A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of image processing technology, specifically relating to a reconstruction model training method, reconstruction method, apparatus, and electronic equipment. Background Technology
[0002] Understanding 3D structures from Multiple View Stereo (MVS) data performs well across multiple tasks, such as dense reconstruction, novel perspective synthesis, and sealed surface reconstruction. Focusing on reconstruction, MVS utilizes images collected from different camera positions to build digital replicas of scenes or objects. While traditional MVS reconstruction workflows perform well in both indoor and outdoor scenes, they still struggle with reconstructing finer details, such as surfaces that are too noisy or incomplete. Summary of the Invention
[0003] The purpose of this application is to provide a reconstruction model training method, reconstruction method, apparatus, and electronic device that can address at least some of the aforementioned problems.
[0004] In a first aspect, embodiments of this application provide a method for training a reconstruction model, including:
[0005] The multi-view images of the sample object are input into the reconstruction model to obtain the three-dimensional model corresponding to the sample object. The multi-view images of the sample object are view images of the sample object from different angles obtained by cameras with different camera poses.
[0006] The target loss value is obtained by comparing the sample object and the 3D model through a target loss function. The target loss function includes a first loss function for optimizing the inconsistency of multi-view images, a second loss function for optimizing low-texture surfaces, and / or a third loss function for optimizing fine-grained structures.
[0007] The reconstruction model is iteratively trained based on the target loss value to obtain an optimized reconstruction model.
[0008] Secondly, embodiments of this application also provide a method for reconstructing a three-dimensional model, including:
[0009] Acquire multi-view images of the target object, wherein the multi-view images of the target object are view images of the target object from different angles obtained by cameras with different camera poses;
[0010] The multi-view image of the target object is input into the optimized reconstruction model to obtain the three-dimensional model of the target object after image rendering.
[0011] Thirdly, embodiments of this application also provide a reconstruction model training apparatus, comprising:
[0012] The first reconstruction module is used to input the multi-view images of the sample object into the reconstruction model to obtain the three-dimensional model corresponding to the sample object. The multi-view images of the sample object are view images of the sample object from different angles obtained by cameras with different camera poses.
[0013] The comparison module is used to compare the sample object and the 3D model through a target loss function to obtain a target loss value. The target loss function includes a first loss function for optimizing the inconsistency of multi-view images, a second loss function for optimizing low-texture surfaces, and / or a third loss function for optimizing fine-grained structures.
[0014] The training module is used to iteratively train the reconstruction model based on the target loss value to obtain an optimized reconstruction model.
[0015] Fourthly, embodiments of this application also provide a three-dimensional model reconstruction apparatus, comprising:
[0016] The acquisition module is used to acquire multi-view images of a target object, wherein the multi-view images of the target object are view images of the target object from different angles obtained by cameras with different camera poses;
[0017] The second reconstruction module is used to input the multi-view image of the target object into the optimized reconstruction model to obtain the three-dimensional model of the target object after image rendering.
[0018] Fifthly, embodiments of this application also provide an electronic device, including a memory, a transceiver, and a processor:
[0019] The memory is used to store computer programs; the transceiver is used to send and receive data under the control of the processor; the processor is used to read the computer programs in the memory and execute the methods described in the first or second aspect above.
[0020] In a sixth aspect, embodiments of this application also provide a processor-readable storage medium storing a computer program for causing a processor to perform the method described in the first or second aspect above.
[0021] In a seventh aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method as described in the first or second aspect.
[0022] In this embodiment, firstly, multi-view images of the sample object are acquired and input into the reconstruction model. The reconstruction model performs 3D reconstruction on the multi-view images of the sample object to obtain the corresponding 3D model. The sample object and the 3D model are compared using a target loss function to obtain a target loss value. The reconstruction model is iteratively trained based on the target loss value to obtain an optimized reconstruction model. If the target loss function includes a first loss function for optimizing inconsistencies between multi-view images, that is, the first loss function can optimize the inconsistencies between multi-view images, making the images from different perspectives more consistent in certain features or representations; if the target loss function includes a second loss function for optimizing low-texture surfaces, for some images with less surface texture information, the second loss function can optimize low-texture surfaces, improving the quality or accuracy of these areas in the processing process, thereby improving the quality of reconstruction or rendering; if the target loss function includes a third loss function for optimizing fine-grained structures, for some fine structures on the sample object, the third loss function can enable the model to better capture and reconstruct these details, making the processed image perform better in these detailed structures. By applying at least one of the three different loss functions mentioned above to the target loss function, the multi-view images are processed comprehensively and meticulously to achieve better processing results and more accurate target loss value calculation, thereby improving the model's image rendering and 3D reconstruction performance. Attached Figure Description
[0023] Figure 1 This is a flowchart illustrating the reconstruction model training method provided in the embodiments of this application;
[0024] Figure 2 This is a flowchart illustrating the method for reconstructing a three-dimensional model provided in an embodiment of this application;
[0025] Figure 3 This is a schematic diagram of the structure of the reconstruction model training device provided in the embodiments of this application;
[0026] Figure 4 This is a schematic diagram of the structure of the three-dimensional model reconstruction device provided in the embodiments of this application;
[0027] Figure 5 This is a structural block diagram of the electronic device provided in the embodiments of this application. Detailed Implementation
[0028] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0029] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0030] The reconstruction model training method provided in this application will be described below with reference to the accompanying drawings, through specific embodiments and application scenarios.
[0031] like Figure 1 As shown in the figure, this application provides a method for training a reconstruction model, which may specifically include the following steps:
[0032] Step 101: Input the multi-view images of the sample object into the reconstruction model to obtain the three-dimensional model corresponding to the sample object. The multi-view images of the sample object are view images of the sample object from different angles obtained by cameras with different camera poses.
[0033] Given a set of calibrated multi-view images, each view image is obtained by taking pictures of a sample object (e.g., an object or a static scene) with a camera in different camera poses (e.g., the camera's position, angle, and orientation in space). In other words, the view image taken by each camera pose represents how the sample object appears from different angles.
[0034] The multi-view images of the sample object are input into the reconstruction model, and the reconstruction model performs three-dimensional reconstruction of the multi-view images of the sample object, thereby obtaining the three-dimensional model of the image rendering corresponding to the sample object.
[0035] Step 102: Compare the sample object and the 3D model using a target loss function to obtain a target loss value. The target loss function includes a first loss function for optimizing inconsistencies in multi-view images, a second loss function for optimizing low-texture surfaces, and / or a third loss function for optimizing fine-grained structures.
[0036] The target loss function compares the sample object and the 3D model to obtain the target loss value. If the target loss function includes a first loss function to optimize inconsistencies between multi-view images (making images from different perspectives more consistent in certain features or representations), it can achieve several advantages. If it includes a second loss function to optimize low-texture surfaces (improving the quality or accuracy of these areas in the processing), it can enhance the quality of reconstruction or rendering. If it includes a third loss function to optimize fine-grained structures (allowing the model to better capture and reconstruct these details), the processed image will perform better in these structural details. By applying at least one of these three different loss functions to the target loss function, the multi-view images are processed comprehensively and meticulously to achieve better processing results and more accurate target loss value calculation.
[0037] The comparison between the sample object and the 3D model is conducted from the same perspective by collecting images of the sample object and the 3D model for comparison, in order to improve the accuracy of the target loss function value.
[0038] In another specific embodiment, the target loss function includes a first loss function, a second loss function, and a third loss function. Thus, by having the above three loss functions with different functions work together on the target loss function, the multi-view image can be further processed in a comprehensive and detailed manner to achieve better processing results and more accurate target loss value calculation.
[0039] Step 103: Iteratively train the reconstruction model based on the target loss value to obtain an optimized reconstruction model.
[0040] The reconstruction model is iteratively trained based on the target loss value to optimize the model parameters and obtain the optimized reconstruction model. Since at least one of the three loss functions with different functions acts on the target loss function, the multi-view image can be processed comprehensively and meticulously through the target loss function, which can improve the image rendering and 3D reconstruction performance of the model.
[0041] The construction process of the first loss function is illustrated below through a specific embodiment:
[0042] The construction process of the first loss function specifically includes steps 1021 to 1022:
[0043] Step 1021: For the target three-dimensional surface point of the sample object, calculate the cross-view blur factor of the first ray of the i-th view image. The first ray is the ray emitted by the camera that took the i-th view image toward the target three-dimensional surface point. The i-th view image is each view image in the multi-view image in turn.
[0044] For typical MVS reconstruction, multiple views Figure 1 Consistency ensures accurate surface reconstruction. However, inconsistent components across multiple views (such as moving objects and specular regions) introduce blurred supervision to the surface. For example, reflective metallic surfaces, since reflection does not constitute the majority of the training images, will generate larger reflection scores if the current rendered color is contaminated by reflection, as the relative difference becomes larger compared to most other surfaces. The Sign Distance Function (SDF) is used to accurately label inconsistent supervision across multiple views, such as specular sub-regions and dynamic objects appearing in static scenes, reducing the weight assigned to reflection supervision. It is expected that reflective surfaces will be considered anomalous and assigned high reflection scores. For 3D surface points located on the first ray of the i-th view image (i.e., target 3D surface points), a cross-view blur factor for the first ray of each view image is calculated. This cross-view blur factor makes SDF regularization more adaptive, improving the accuracy of 3D reconstructed surfaces.
[0045] It should be noted that the i-th view image refers to each view image in the multi-view image set sequentially. For example, if the number of multi-view images is 3, i takes values of 1, 2, and 3 sequentially. For the target 3D surface point, the cross-view blur factor of the first ray of the first view image is calculated, as are the cross-view blur factors of the first ray of the second view image and the first ray of the third view image. Here, the first ray of the first view image represents the ray emitted from the camera capturing the first view image towards the target 3D surface point; the first ray of the second view image represents the ray emitted from the camera capturing the second view image towards the target 3D surface point; and the first ray of the third view image represents the ray emitted from the camera capturing the second view image towards the target 3D surface point.
[0046] Step 1022: Construct the first loss function based on the multi-view image, the cross-view blur factor of the first ray of the i-th view image, the rendered color of the first ray, and the real color of the first ray.
[0047] Based on multi-view images, the cross-view blur factor of the first ray of each view image, the rendered color of the first ray of each view image, and the true color of the first ray of each view image, a first loss function is constructed to optimize the inconsistency between multi-view images. The first loss function makes the images under different perspectives more coordinated and unified in certain features or performance.
[0048] In another optional specific embodiment, step 1021, for the target three-dimensional surface points of the sample object, calculates the cross-view blur factor of the first ray of the i-th view image, specifically including:
[0049] For the target three-dimensional surface point of the sample object, obtain the rendering color of the first ray of the i-th view image and the rendering color of the first ray of the q-th view image. The first ray of the q-th view image is the ray emitted by the camera that captures the q-th view image toward the target three-dimensional surface point. The q-th view image is each view image in the multi-view image in turn.
[0050] Based on the rendered color of the first ray of the i-th view image and the rendered color of the first ray of the q-th view image, calculate the cross-view blur factor of the first ray of the i-th view image.
[0051] By enforcing radiative supervision through multi-view images, Neural Radiance Fields (NeRF) utilizes volumetric rendering to match the ground truth (i.e., true color) of each camera pose with the rendered image (i.e., rendered color). Specifically, red, green, and blue (RGB) values can be generated for each pixel of the image by sampling n points along the first ray of its camera, as expressed by the following formula:
[0052] r(t j )=o+t j *v
[0053] Where: o represents the center position of the camera;
[0054] t j This represents the sampling interval along the first ray, or it can represent the sampling point t. j j takes values from 1 to n.
[0055] v indicates the view direction.
[0056] The rendering color of the first ray can be calculated by accumulating the radiation field density and rendering color of the sampled points.
[0057] The rendering color of the first ray of each view image is obtained in the manner described above. The q-th view image is each view image in the multi-view image in turn. For example, if the number of multi-view images is 3, then q takes values of 1, 2, and 3 in turn. For the target 3D surface point, the rendering color of the first ray of the first view image is obtained, as are the rendering colors of the first ray of the second view image and the first ray of the third view image.
[0058] The cross-view blur factor of the first ray of the i-th view image is calculated as follows: based on the color difference between the rendered color of the first ray of the i-th view image and the rendered color of the first ray of each view image, the cross-view blur factor of the first ray of the i-th view image is calculated.
[0059] In one specific embodiment, the cross-view blur factor of the first ray of the i-th view image is specifically calculated using the following formula:
[0060]
[0061] Where m represents the number of view images;
[0062] r i Represents the first ray of the i-th view image;
[0063] C(r i ) represents the rendered color of the first ray of the i-th view image;
[0064] r q Represents the first ray of the q-th view image;
[0065] C(r q ) represents the rendered color of the first ray of the q-th view image;
[0066] (*) T This represents the transpose of (*).
[0067] ∑ -1 (*) denotes the inverse of the empirical covariance matrix;
[0068] V q The visibility verification factor represents the first ray of the q-th view image;
[0069] λ c (r i ) represents the cross-view blur factor of the first ray of the i-th view image.
[0070] Visibility verification factor V qEach ray is labeled to avoid misleading information introduced by invisible rays. Calculating the visibility verification factor by comparing the intersection of the ray and the intermediate reconstructed mesh through ray projection is too sensitive to the resolution of the skinbox mesh production stage, leading to false visibility alarms when the ray hits smaller structures. To address this, the visibility verification factor is calculated directly along the ray in the SDF by sampling points. The activations of these sampling points are learned by the ProposalNet network, and this activation is parameterized using a small MultiLayer Perceptron (MLP). This is done because the mechanism collects more true positive points in occluded areas, contributing to accurate target detection. To reduce the number of false alarms, these points are still filtered by the SDF. When adjacent sampling points have different signs, bilinear interpolation is used to determine and record surface points along the ray direction using the corresponding SDF values of these two points. Bilinear interpolation is a method of interpolating data in two-dimensional space, used here to more accurately locate surface points.
[0071] It should be noted that the visibility verification factor V of the first ray in the q-th view image q A value of 0 indicates that the 3D surface points of the first ray in the q-th view image are occluded. The visibility verification factor V of the first ray in the q-th view image. q When the value is 1, it indicates that the 3D surface point of the first ray of the q-th view image is not occluded.
[0072] In one specific embodiment, the first loss function is specifically represented by the following formula:
[0073]
[0074] Where m represents the number of view images;
[0075] r i Represents the first ray of the i-th view image;
[0076] Represents the rendered color of the first ray of the i-th view image;
[0077] C(r i ) represents the true color of the first ray in the i-th view image;
[0078] λ c (r i ) represents the cross-view blur factor of the first ray of the i-th view image;
[0079] ‖*‖2 represents the Euclidean distance formula;
[0080] |*| represents the L1 Euclidean distance formula;
[0081] L rgb This represents the first loss function.
[0082] Since the image rendering process is differentiable, the model can then learn the radiation field from the multi-view image. The first loss function minimizes the color difference between the rendered color and the corresponding real ground color. For example, the model reconstructed by the first loss function does not miss the small metal handle on the side and there is no collapse phenomenon, which improves the accuracy of the 3D reconstructed surface.
[0083] The construction process of the second loss function is illustrated below through a specific example:
[0084] The construction process of the second loss function specifically includes steps 1023 to 1025:
[0085] Step 1023: For the first ray of the i-th view image, obtain the symbolic distance function (SDF value) value of each sampling point of the first ray.
[0086] Based on proper multi-view stereo (MVS) supervision, the learned set {S,C} is determined by the signed distance function f(x):R 3 →R represents a radiant field c(x,v) where the value of each element is determined by its 3D position (x, y, z) and a radiant field c(x,v):R 3 ×S 2 →R 3 It consists of position (x, y, z) and viewing direction v∈S 2 A joint decision.
[0087] For the first ray of the i-th view image, the SDF value of each sampling point of the first ray of the i-th view image is obtained in the manner described above.
[0088] Step 1024: Based on the symbol distance function value of each sampling point of the first ray, calculate the local radiation difference factor of each sampling point of the first ray and the finest-grained patch term of the first ray.
[0089] Learning neural surfaces presents challenges in accurately reconstructing low-texture regions because per-pixel rendering is readily satisfied compared to ground truth. Aside from the boundaries of low-texture surfaces, few effective gradients are backpropagated to correct erroneously learned SDF. Let r(t0) be the point on the first ray closest to the zero-crossing point. A flat local surface should possess similar properties: the radiance observed from the sampling point onto the sample object surface and the SDF value that varies linearly along the ray, in order to improve the accuracy of continuity and smoothness in the region by using the ray itself.
[0090] To correctly enforce these two properties as regularization during training, a second loss function with the finest-grained patch term and planar regularization is used to optimize low-texture surfaces. Furthermore, to learn more accurate zero-crossing surfaces by jointly training the SDF and radiation field, the SDF regularization is made more adaptive through a local radiometric difference factor, improving both the smoothness of low-texture regions and the sharpness of fine-grained structures.
[0091] Step 1025: Construct the second loss function based on the symbolic distance function value and local radiation difference factor of each sampling point of the first ray, and the finest-grained patch term of the first ray.
[0092] Based on the SDF value of each sampling point of the first ray, the local radiometric difference factor of each sampling point of the first ray, and the finest-grained patch term of the first ray, a second loss function is constructed to optimize low-texture surfaces. For some images with less surface texture information, the second loss function can optimize low-texture surfaces, improve the quality or accuracy of these regions in the processing, and thus improve the quality of reconstruction or rendering.
[0093] In one specific embodiment, the finest-grained patch item of the first ray is calculated using the following formula:
[0094]
[0095] r(t0) represents the position of the sampling point t0 of the first ray;
[0096] r(t1) represents the position of the sampling point t1 of the first ray;
[0097] f(r(t1)) represents the SDF value of the sampling point t1 of the first ray;
[0098] ‖*‖2 represents the Euclidean distance formula;
[0099] P represents the finest-grained patch item for the first ray.
[0100] The finest-grained patch term represents the ratio of the SDF difference f(r(t1))-f(r(t0)) to the Euclidean distance ‖r(t1)-r(t0)‖2 of these sampling points, where f(t(t0)) is actually 0.
[0101] In one specific embodiment, the second loss function is specifically represented by the following formula:
[0102]
[0103] Where n represents the number of sampling points for the first ray;
[0104] r(tj ) represents the sampling point t of the first ray. j Location;
[0105] f(r(t j )) represents the sampling point t of the first ray. j SDF value;
[0106] r(t0) represents the position of the sampling point t0 of the first ray;
[0107] The sampling point t represents the first ray. j Local radiation difference factor;
[0108] P represents the finest-grained patch item for the first ray;
[0109] ‖*‖2 represents the Euclidean distance formula;
[0110] L planar This represents the second loss function.
[0111] When the surrounding surface is an ideal plane, the second loss function should be fully executed. To automatically determine whether a local surface is close to an ideal plane, the second loss function can be reweighted using a local radiation difference factor. Intuitively, when viewed from a direction perpendicular to the surface, and disregarding other slight blurring caused by shadows and different lighting intensities, flat surface patches should have a very similar appearance.
[0112] In one specific embodiment, the local radiation difference factor at each sampling point of the first ray is calculated using the following formula:
[0113]
[0114] Wherein, C(r) j ′) represents the sampling point t from the first ray. j The rendered color observed from the point on the target 3D surface at the location of the point sample t. j The radiation observed at the location of the zero-crossing surface point in the SDF;
[0115] C(r1) represents the rendered color observed from the sampling point t1 of the first ray to the target 3D surface point, that is, the radiation observed from the sampling point t1 to the zero-crossing surface point in the SDF;
[0116] ‖*‖2 represents the Euclidean distance formula;
[0117] The sampling point t represents the first ray. j The local radiation difference factor.
[0118] Through the radiation difference ||C(r′) j )-C(r1)||2 reweighted each Among the larger ones This indicates that the subregion is more likely to be flat. C(r) j ′) was calculated using the new rays:
[0119]
[0120] Among them, the new surface orientation view direction It is calculated using the gradient direction of SDF:
[0121]
[0122] Among them, the approximate normal line It is f(r(t) j The derivative of ).
[0123] The construction process of the third loss function is illustrated below through a specific example:
[0124] The construction process of the third loss function specifically includes steps 1026 to 1028:
[0125] Step 1026: For the target three-dimensional surface points of the sample object, obtain the signed distance function value of each sampling point of the first ray of the i-th view image.
[0126] Based on proper multi-view stereo (MVS) supervision, the learned set {S,C} is determined by the signed distance function f(x):R 3 →R represents a radiant field c(x,v) where the value of each element is determined by its 3D position (x, y, z) and a radiant field c(x,v):R 3 ×S 2 →R 3 It consists of position (x, y, z) and viewing direction v∈S 2 A joint decision.
[0127] For the first ray of the i-th view image, the SDF value of each sampling point of the first ray of the i-th view image is obtained in the manner described above.
[0128] Step 1027: Based on the symbolic distance function value of each sampling point of the first ray of the i-th view image, calculate the local radiometric difference factor and the derivative of the symbolic distance function value of each sampling point of the first ray of the i-th view image.
[0129] To learn more accurate zero-crossing surfaces by jointly training the SDF and radiation field, the SDF regularization is made more adaptive by using a local radiation difference factor, which improves both the smoothness of low-texture regions and the sharpness of fine-grained structures.
[0130] Based on the SDF value of each sampling point of the first ray of the i-th view image, calculate the local radiometric difference factor of each sampling point of the first ray of the i-th view image and the derivative of the SDF value of each sampling point of the first ray of the i-th view image.
[0131] Step 1028: Construct the third loss function based on the local radiation difference factor of each sampling point of the first ray of the i-th view image and the derivative of the symbolic distance function value.
[0132] To address the issue of failed fine-structure reconstruction and prevent the SDF from including the corresponding structures, the ability to render such structures from the radiation field should be ensured. This allows for subsequent backpropagation of gradients from the radiation field to the SDF, and optimization of the SDF using a third loss function. Based on the local radiometric difference factor at each sampling point of the first ray and the derivative of the SDF value at each sampling point of the first ray, a third loss function is constructed to optimize fine-grained structures. For subtle structures on the sample object, the third loss function enables the model to better capture and reconstruct these details, resulting in better performance of the processed image in these detailed structures.
[0133] In one specific embodiment, the third loss function is specifically represented by the following formula:
[0134]
[0135] Where, λ E This represents the constant 0.1;
[0136] m represents the number of view images;
[0137] n represents the number of sampling points for the first ray of the i-th view image;
[0138] n ij This represents the derivative of the SDF value at the j-th sampling point of the first ray in the i-th view image;
[0139] The sampling point t represents the first ray. j Local radiation difference factor;
[0140] ‖*‖2 represents the Euclidean distance formula;
[0141] L sdf This represents the third loss function.
[0142] Through the radiation difference ||C(r′) j )-C(r1)||2 reweighted each Among the larger ones The local pattern is complex, therefore SDF regularization should be relaxed to prioritize rendering. Thus, improving rendering quality should be prioritized over forcing Eikonal regularization to achieve a more perfect SDF. In the specific implementation, approximate normals are used. It is f(r(t) j The derivative of )), λ E Typically set to 0.1. Adaptive Eikonal regularization addresses the lack of detailed structure in methods that only address reflection blur, revealing the true structure. This is achieved through simple application... It can oversmooth the correct geometry, revealing more details without causing the reflective water surface to collapse.
[0143] In one specific embodiment, the target loss function is the sum of the first loss function, the second loss function, and / or the third loss function.
[0144] For example, the target loss function is specifically expressed by the following formula:
[0145] L total =L rgb +L planar +L sdf
[0146] Among them, L rgb Represents the first loss function;
[0147] L planar Represents the second loss function
[0148] L sdf Represents the third loss function;
[0149] L total This represents the target loss function.
[0150] The target loss function is obtained by adding the first, second, and third loss functions mentioned above. Adaptive weighted Eikonal regularization is used to optimize the SDF. By having three loss functions with different functions work together on the target loss function, the multi-view image can be processed in a comprehensive and detailed manner to achieve better processing results and more accurate target loss value calculation.
[0151] The reconstruction model was trained to simultaneously perform geometry optimization in 3D space and appearance optimization in 2D space, with an objective loss function of L. total The Adam optimizer was employed, with an exponentially decaying learning rate schedule ranging from 1*10⁻² to 1*10⁻⁵.
[0152] The model is trained using the cross-view blur factor proposed in the first loss function, and the local radiance difference factor proposed in the second and third loss functions. It can be built on top of NeuS2 (a neural surface reconstruction method for multi-view reconstruction) and a customized background model. Progressive training is applied during training to stabilize the low-level hash table, where the first four levels are optimized together in the first 2000 steps, and a new higher level is included for optimization every 2000 steps. To deploy the model in unbounded scenes, the foreground and background spaces are represented in two separate models. The foreground space is modeled by NeuS2 within a cubic bounding box, covering the near-field scene; while the background is modeled separately by a vanilla NeRF (Neural Radiance Fields) model, encoded by an MLP without SDF middleware, which learns in a contracted spherical space with a radius twice the length of the cube's diagonal.
[0153] In summary, the embodiments of this application include a first loss function for optimizing inconsistencies in multi-view images. This first loss function optimizes inconsistencies between multi-view images, making images from different perspectives more consistent in certain features or representations. The target loss function also includes a second loss function for optimizing low-texture surfaces. For images with limited surface texture information, the second loss function optimizes low-texture surfaces, improving the quality or accuracy of these areas during processing, thereby enhancing the quality of reconstruction or rendering. Finally, the target loss function includes a third loss function for optimizing fine-grained structures. For subtle structures on sample objects, the third loss function enables the model to better capture and reconstruct these details, resulting in better performance of the processed image in these detailed structures. By applying at least one of these three loss functions to the target loss function, inconsistencies in multi-view radiation are mitigated, the depiction of low-texture surfaces is strengthened, and the complexity of fine-grained structures is learned within a unified framework. This allows for comprehensive and detailed processing of multi-view images, achieving better processing results and more accurate target loss value calculation, thereby improving the model's image rendering and 3D reconstruction performance.
[0154] like Figure 2 As shown in the embodiments of this application, a method for reconstructing a three-dimensional model is also provided, specifically including:
[0155] Step 201: Obtain multi-view images of the target object. The multi-view images of the target object are view images of the target object taken by cameras with different camera poses from different angles.
[0156] By using a camera to capture images of a target object (such as an object or a static scene) in different camera poses (e.g., the camera's position, angle, and orientation in space), multi-view images are obtained. In other words, the view image captured by each camera pose represents how the target object appears from different angles.
[0157] Step 202: Input the multi-view image of the target object into the optimized reconstruction model to obtain the three-dimensional model of the target object after image rendering.
[0158] The multi-view images of the target object are input into the optimized reconstruction model. The optimized reconstruction model is used to perform three-dimensional reconstruction of the multi-view images of the target object, thereby obtaining the three-dimensional model of the target object after image rendering.
[0159] Because the optimized reconstruction model can reduce inconsistencies in multi-view radiation, enhance the depiction of low-texture surfaces, and perform comprehensive and detailed processing of multi-view images, it can improve image rendering and 3D reconstruction performance, resulting in a more accurate 3D model.
[0160] The reconstruction model training method provided in the above embodiments can be executed by a reconstruction model training device. This application uses the reconstruction model training device executing the reconstruction model training method as an example to illustrate the reconstruction model training device provided in this application embodiment.
[0161] like Figure 3 As shown in the illustration, this application also provides a reconstruction model training device 300, which specifically includes:
[0162] The first reconstruction module 301 is used to input the multi-view images of the sample object into the reconstruction model to obtain the three-dimensional model corresponding to the sample object. The multi-view images of the sample object are view images of the sample object from different angles obtained by cameras with different camera poses.
[0163] The comparison module 302 is used to compare the sample object and the three-dimensional model through a target loss function to obtain a target loss value. The target loss function includes a first loss function for optimizing the inconsistency of multi-view images, a second loss function for optimizing low-texture surfaces, and / or a third loss function for optimizing fine-grained structures.
[0164] The training module 303 is used to iteratively train the reconstruction model based on the target loss value to obtain an optimized reconstruction model.
[0165] Optionally, the apparatus further includes a first loss function construction module, specifically used for:
[0166] For the target three-dimensional surface point of the sample object, calculate the cross-view blur factor of the first ray of the i-th view image, where the first ray is the ray emitted by the camera that captured the i-th view image toward the target three-dimensional surface point, and the i-th view image is each view image in the multi-view image in turn.
[0167] The first loss function is constructed based on the multi-view image, the cross-view blur factor of the first ray of the i-th view image, the rendered color of the first ray, and the real color of the first ray.
[0168] Optionally, when the first loss function construction module calculates the cross-view blur factor of the first ray of the i-th view image for the target 3D surface points of the sample object, it is specifically used for:
[0169] For the target three-dimensional surface point of the sample object, obtain the rendering color of the first ray of the i-th view image and the rendering color of the first ray of the q-th view image. The first ray of the q-th view image is the ray emitted by the camera that captures the q-th view image toward the target three-dimensional surface point. The q-th view image is each view image in the multi-view image in turn.
[0170] Based on the rendered color of the first ray of the i-th view image and the rendered color of the first ray of the q-th view image, calculate the cross-view blur factor of the first ray of the i-th view image.
[0171] Optionally, the apparatus further includes a second loss function construction module, specifically used for
[0172] For the first ray of the i-th view image, obtain the signed distance function value of each sampling point of the first ray;
[0173] Based on the symbol distance function value of each sampling point of the first ray, calculate the local radiation difference factor of each sampling point of the first ray and the finest-grained patch term of the first ray;
[0174] The second loss function is constructed based on the symbolic distance function value and local radiation difference factor of each sampling point of the first ray, and the finest-grained patch term of the first ray.
[0175] Optionally, the apparatus further includes a third loss function construction module, specifically used for
[0176] For the target three-dimensional surface points of the sample object, obtain the signed distance function value of each sampling point of the first ray of the i-th view image;
[0177] Based on the symbolic distance function value of each sampling point of the first ray of the i-th view image, calculate the local radiometric difference factor and the derivative of the symbolic distance function value of each sampling point of the first ray of the i-th view image;
[0178] The third loss function is constructed based on the local radiometric difference factor of each sampling point of the first ray of the i-th view image and the derivative of the sign distance function value.
[0179] Optionally, the target loss function is the sum of the first loss function, the second loss function, and / or the third loss function.
[0180] It should be noted that the reconstruction model training device provided in this application embodiment can implement all the method steps implemented in the above reconstruction model training method embodiment and can achieve the same technical effect. Here, the parts that are the same as those in the method embodiment and the beneficial effects will not be described in detail.
[0181] The three-dimensional model reconstruction method provided in the above embodiments can be executed by a three-dimensional model reconstruction device. This application uses a three-dimensional model reconstruction device executing the three-dimensional model reconstruction method as an example to illustrate the three-dimensional model reconstruction device provided in this application embodiment.
[0182] like Figure 4 As shown in the figure, this application embodiment also provides a three-dimensional model reconstruction device 400, specifically including:
[0183] The acquisition module 401 is used to acquire multi-view images of a target object, wherein the multi-view images of the target object are view images of the target object taken by cameras with different camera poses from different angles;
[0184] The second reconstruction module 402 is used to input the multi-view image of the target object into the optimized reconstruction model to obtain the three-dimensional model of the target object after image rendering.
[0185] It should be noted that the three-dimensional model reconstruction apparatus provided in this application embodiment can implement all the method steps implemented in the three-dimensional model reconstruction method embodiment and can achieve the same technical effect. Here, the parts that are the same as those in the method embodiment and the beneficial effects will not be described in detail.
[0186] It should be noted that the division of units in the embodiments of this application is illustrative and only represents one logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units described above can be implemented in hardware or as software functional units.
[0187] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a processor-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0188] like Figure 5 As shown, embodiments of this application also provide an electronic device, including a memory 520, a transceiver 510, and a processor 500:
[0189] Memory 520 is used to store computer programs;
[0190] Transceiver 510 is used to send and receive data under the control of the processor;
[0191] The processor 500 is configured to read a computer program from a memory and execute the steps of the reconstruction model training method as described in any of the above embodiments, or execute the steps of the three-dimensional model reconstruction method as described in any of the above embodiments.
[0192] Among them, Figure 5In this context, the bus architecture can include any number of interconnected buses and bridges, specifically linking various circuits together, represented by one or more processors (processor 500) and memory (memory 520). The bus architecture can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. The bus interface provides an interface. The transceiver 510 can be multiple elements, including transmitters and receivers, providing a unit for communicating with various other devices over transmission media, including wireless channels, wired channels, optical fibers, etc. The processor 500 is responsible for managing the bus architecture and general processing, and the memory 520 can store data used by the processor 500 during operation.
[0193] The processor 500 can be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a complex programmable logic device (CPLD). The processor can also adopt a multi-core architecture.
[0194] The processor executes the reconstruction model training method provided in this application embodiment by calling a computer program stored in memory, according to the obtained executable instructions. The processor and memory can also be physically separated.
[0195] It should be noted that the electronic device provided in this application embodiment can implement all the method steps implemented in the above reconstruction model training method or 3D model reconstruction method embodiment, and can achieve the same technical effect. Here, the parts that are the same as those in the method embodiment and the beneficial effects will not be described in detail.
[0196] Embodiments of this application also provide a processor-readable storage medium storing a computer program for causing the processor to execute the above-described reconstruction model training method or 3D model reconstruction method.
[0197] This application also provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described reconstruction model training method or the above-described three-dimensional model reconstruction method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0198] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0199] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) to execute the methods described in the various embodiments of this application.
[0200] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A method for training a reconstruction model, characterized in that, include: The multi-view images of the sample object are input into the reconstruction model to obtain the three-dimensional model corresponding to the sample object. The multi-view images of the sample object are view images of the sample object from different angles obtained by cameras with different camera poses. The target loss value is obtained by comparing the sample object and the 3D model through a target loss function. The target loss function includes a first loss function for optimizing the inconsistency of multi-view images, a second loss function for optimizing low-texture surfaces, and / or a third loss function for optimizing fine-grained structures. The reconstruction model is iteratively trained based on the target loss value to obtain an optimized reconstruction model.
2. The method according to claim 1, characterized in that, The process of constructing the first loss function includes: For the target three-dimensional surface point of the sample object, calculate the cross-view blur factor of the first ray of the i-th view image, where the first ray is the ray emitted by the camera that captured the i-th view image toward the target three-dimensional surface point, and the i-th view image is each view image in the multi-view image in turn. The first loss function is constructed based on the multi-view image, the cross-view blur factor of the first ray of the i-th view image, the rendered color of the first ray, and the real color of the first ray.
3. The method according to claim 2, characterized in that, The calculation of the cross-view blur factor of the first ray of the i-th view image for the target three-dimensional surface points of the sample object includes: For the target three-dimensional surface point of the sample object, obtain the rendering color of the first ray of the i-th view image and the rendering color of the first ray of the q-th view image. The first ray of the q-th view image is the ray emitted by the camera that captures the q-th view image toward the target three-dimensional surface point. The q-th view image is each view image in the multi-view image in turn. Based on the rendered color of the first ray of the i-th view image and the rendered color of the first ray of the q-th view image, calculate the cross-view blur factor of the first ray of the i-th view image.
4. The method according to any one of claims 2 to 3, characterized in that, The process of constructing the second loss function includes: For the first ray of the i-th view image, obtain the signed distance function value of each sampling point of the first ray; Based on the symbol distance function value of each sampling point of the first ray, calculate the local radiation difference factor of each sampling point of the first ray and the finest-grained patch term of the first ray; The second loss function is constructed based on the symbolic distance function value and local radiation difference factor of each sampling point of the first ray, and the finest-grained patch term of the first ray.
5. The method according to any one of claims 2 to 4, characterized in that, The process of constructing the third loss function includes: For the target three-dimensional surface points of the sample object, obtain the signed distance function value of each sampling point of the first ray of the i-th view image; Based on the symbolic distance function value of each sampling point of the first ray of the i-th view image, calculate the local radiometric difference factor and the derivative of the symbolic distance function value of each sampling point of the first ray of the i-th view image; The third loss function is constructed based on the local radiometric difference factor of each sampling point of the first ray of the i-th view image and the derivative of the sign distance function value.
6. The method according to any one of claims 1 to 5, characterized in that, The target loss function is the sum of the first loss function, the second loss function, and / or the third loss function.
7. A method for reconstructing a three-dimensional model, characterized in that, include: Acquire multi-view images of the target object, wherein the multi-view images of the target object are view images of the target object from different angles obtained by cameras with different camera poses; The multi-view image of the target object is input into the optimized reconstruction model to obtain the three-dimensional model of the target object after image rendering.
8. A reconstruction model training device, characterized in that, include: The first reconstruction module is used to input the multi-view images of the sample object into the reconstruction model to obtain the three-dimensional model corresponding to the sample object. The multi-view images of the sample object are view images of the sample object from different angles obtained by cameras with different camera poses. The comparison module is used to compare the sample object and the 3D model through a target loss function to obtain a target loss value. The target loss function includes a first loss function for optimizing the inconsistency of multi-view images, a second loss function for optimizing low-texture surfaces, and / or a third loss function for optimizing fine-grained structures. The training module is used to iteratively train the reconstruction model based on the target loss value to obtain an optimized reconstruction model.
9. A device for reconstructing a three-dimensional model, characterized in that, include: The acquisition module is used to acquire multi-view images of a target object, wherein the multi-view images of the target object are view images of the target object from different angles obtained by cameras with different camera poses; The second reconstruction module is used to input the multi-view image of the target object into the optimized reconstruction model to obtain the three-dimensional model of the target object after image rendering.
10. An electronic device, characterized in that, Includes memory, transceiver, and processor: The memory is used to store computer programs; the transceiver is used to send and receive data under the control of the processor; the processor is used to read the computer programs in the memory and execute the steps of the reconstruction model training method as described in any one of claims 1-6, or execute the steps of the three-dimensional model reconstruction method as described in claim 7.
11. A processor-readable storage medium, characterized in that, The processor-readable storage medium stores a computer program that causes the processor to perform the steps of the reconstruction model training method as described in any one of claims 1-6, or the steps of the three-dimensional model reconstruction method as described in claim 7.
12. A computer program product, characterized in that, The computer program product is stored in a storage medium and is executed by at least one processor to implement the steps of the reconstruction model training method as described in any one of claims 1-6, or to implement the steps of the three-dimensional model reconstruction method as described in claim 7.